Text processing method and device
By grouping text sequences into text groups with high semantic similarity, the problem of input length limitation of large language models in long text scenarios is solved, and more efficient text processing and model output is achieved.
Patent Information
- Application Number
- CN202311734367.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-17
AI Technical Summary
The existing large language model is not effective in long text scenarios, mainly due to the limitation of input text length and the secondary increase in inference costs, resulting in limited application in books, academic papers, legal contracts and other scenarios.
By grouping text sequences into multiple text groups, the semantic similarity of subtext sequences within each group is greater than the threshold, this grouping method improves the effective input length of the language model while maintaining processing accuracy.
It effectively expands the input length of the language model, reduces the loss of text semantic information, and improves the compression effect and the accuracy of model output.
Smart Images

Figure CN120163131A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a text processing method and apparatus thereof. Background Art
[0002] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0003] Large language models are neural network models with a large number of parameters trained on a large amount of text data. They can accurately understand the meaning of language texts or generate specified natural language texts, and effectively handle various natural language processing tasks, such as important tasks like translation, dialogue, and code generation. Therefore, they have great commercial value. However, current large language models often have limitations on the length of the input text, making it impossible to effectively deploy them in long text scenarios, such as books, academic papers, legal contracts, etc.
[0004] The limitations on the length of the input text of language models usually have two main reasons. One is that language models are often trained on short texts, and the positional encoding of large language models does not directly adapt to the situation of long texts, resulting in a significant decline in the performance of large language models on long texts. The other is that the inference cost of language models grows quadratically with the length of the input text, and the longer the text, the higher the load on the inference engine and platform.
[0005] Therefore, there is an urgent need for a method that can increase the effective input length of language models. Summary of the Invention
[0006] This application provides a text processing method, which groups multiple sub-text sequences in a text sequence and performs subsequent processing of the language model at the granularity of text groups. Since the semantic proximity between the sub-text sequences within a text group, it can ensure the processing accuracy of the language model while increasing the effective input length of the language model.
[0007] In a first aspect, the present application provides a text processing method, the method comprising: obtaining a text sequence; the text sequence comprising a plurality of sub-text sequences; grouping the plurality of sub-text sequences to obtain a plurality of text groups; the semantic similarity between the sub-text sequences comprised in the same text group being greater than a first threshold; the plurality of text groups being used to obtain a model output through a language model.
[0008] In the present application, the plurality of sub-text sequences in the text sequence are grouped. Each group comprises sub-text sequences with a reduced length relative to the length of the text sequence. And because the sub-text sequences within the text group are semantically close, that is to say, the sub-text sequences within the text group are continuous (a text group composed of semantically close sub-text sequences can express a common or similar theme. In this case, the continuity and coherence of the sub-text sequences within the text group are higher). The continuous sub-text sequences can enable the text groups obtained after grouping to not lose too much semantic information. When performing subsequent processing of the language model with the text groups as the granularity, the accuracy of the model output can be guaranteed. Furthermore, when the effective input length of the language model is increased, the processing accuracy of the language model is guaranteed.
[0009] Among them, the text sequence is one of the common serialized data types. The text sequence can also be referred to as text data. The text data can be regarded as a sequence of characters or a sequence of words, that is, the text sequence can be data obtained by arranging sub-words or words in a certain order. Each sub-text sequence can be a continuous partial sequence of the text sequence.
[0010] Among them, after obtaining the plurality of text groups, other processing such as compressing and sequence expanding each text group can also be performed.
[0011] In a possible implementation, the method further comprises: compressing each text group among the plurality of text groups to obtain a plurality of processing results; the plurality of text groups being used to obtain a model output through a language model, including: the plurality of processing results being used to obtain a model output through a language model.
[0012] The present application compresses a long text sequence input to the language model (when the text sequence exceeds the maximum input length supported by the language model, the text sequence can be called a long text) into a shorter natural language text, thus equivalently expanding the input length of the large language model. In addition, in the embodiments of the present application, the sub-text sequences within each text group obtained after grouping describe similar semantic content, which can maintain the coherence of the content of each sub-text sequence (the semantics of text content with higher continuity and coherence are similar), thereby reducing the loss of text semantic information during compression and improving the compression effect.
[0013] In a possible implementation, the semantic similarity between the sub-text sequences included in different text groups is less than the first threshold.
[0014] To ensure a relatively high continuity between the sub-text sequences within the same text group (that is, their positions in the original text sequence are consecutive or very close, so that the compression process can identify richer continuous semantic information, thereby improving the compression quality), when grouping the sub-text sequences, in addition to considering semantic similarity, the position similarity between the sub-text sequences in the text sequence can also be considered, that is, by grouping, the position similarity between the sub-text sequences included in the same text group in the text sequence is greater than the second threshold. In addition, it can also be achieved by grouping that the position similarity between the sub-text sequences included in the same text group in the text sequence is less than the second threshold.
[0015] In a possible implementation, there is an overlap between at least two of the multiple sub-text sequences. The existence of overlap can improve the continuity between the sub-text sequences.
[0016] In a possible implementation, the method further includes: splicing the multiple processing results according to the positional relationship between the multiple text groups in the text sequence to obtain the spliced processing result.
[0017] In a possible implementation, the method further includes: obtaining the processing result of the text sequence through a language model according to the multiple text groups.
[0018] In a possible implementation, the step of compressing each of the multiple text groups to obtain multiple processing results includes: compressing each of the multiple text groups through the language model or other machine learning models other than the language model to obtain multiple processing results.
[0019] Machine learning for compression also has limitations on the length of the input text. Through the method of the present application, the text sequence is grouped, and the long text input to the compression model is segmented and grouped into shorter text groups during compression. This equivalently extends the input length of the compression model, and each text group includes sub-text sequences describing similar semantic content, maintaining the coherence of the content of each sub-text sequence, thereby reducing the loss of text semantic information during compression and improving the compression effect.
[0020] In a possible implementation, the model output is the result of performing one of the following tasks on the text sequence: abstract extraction task, text generation task, dialogue task, question and answer task, text translation task, knowledge retrieval task.
[0021] Second aspect, the present application provides a text processing device, the device comprising:
[0022] An acquisition module, configured to acquire a text sequence; the text sequence includes a plurality of sub - text sequences;
[0023] A grouping module, configured to group the plurality of sub - text sequences to obtain a plurality of text groups; the semantic similarity between the sub - text sequences included in the same text group is greater than a first threshold; the plurality of text groups are used to obtain a model output through a language model.
[0024] In a possible implementation, the device further comprises:
[0025] A compression module, configured to compress each text group in the plurality of text groups to obtain a plurality of processing results;
[0026] The plurality of text groups being used to obtain a model output through a language model includes:
[0027] The plurality of processing results are used to obtain a model output through a language model.
[0028] In a possible implementation, the semantic similarity between the sub - text sequences included in different text groups is less than the first threshold.
[0029] In a possible implementation, the position similarity between the sub - text sequences included in the same text group in the text sequence is greater than a second threshold.
[0030] In a possible implementation, the position similarity between the sub - text sequences included in the same text group in the text sequence is less than the second threshold.
[0031] In a possible implementation, there is an overlap between at least two sub - text sequences among the plurality of sub - text sequences.
[0032] In a possible implementation, the compression module is further configured to:
[0033] According to the position relationship between the plurality of text groups in the text sequence, splice the plurality of processing results to obtain a spliced processing result.
[0034] In a possible implementation, the device further comprises:
[0035] A task processing module, configured to obtain a processing result of the text sequence through a language model according to the plurality of text groups.
[0036] In a possible implementation, the compression module is specifically configured to:
[0037] Compress each of the multiple text groups through the language model or other machine learning models other than the language model to obtain multiple processing results.
[0038] In a possible implementation, the model output is the result of performing one of the following tasks on the text sequence:
[0039] Abstract extraction task, text generation task, dialogue task, question and answer task, text translation task, knowledge retrieval task.
[0040] In a third aspect, an embodiment of the present application provides a text processing device, which may include a memory, a processor, and a bus system. Among them, the memory is used to store programs, and the processor is used to execute the programs in the memory to execute the methods in the first aspect and any optional method thereof as described above.
[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it causes the computer to execute the methods in the first aspect and any optional method thereof as described above.
[0042] In a fifth aspect, an embodiment of the present application provides a computer program, which, when running on a computer, causes the computer to execute the methods in the first aspect and any optional method thereof as described above.
[0043] In a sixth aspect, the present application provides a chip system, which includes a processor for supporting the execution of the functions involved in the above aspects by a text processing device. For example, sending or processing the data or information involved in the above methods. In a possible design, the chip system further includes a memory for storing necessary program instructions and data for an execution device or a training device. The chip system may be composed of chips or may include chips and other discrete devices. Description of the Drawings
[0044] Figure 1A It is a schematic structural diagram of a framework of an artificial intelligence entity;
[0045] Figure 1B and to Figure 1C It is a schematic diagram of an application system framework of the present application;
[0046] Figure 1D It is a schematic diagram of an optional hardware structure of a terminal;
[0047] Figure 2 It is a schematic diagram of the structure of a server;
[0048] Figure 3 It is a schematic diagram of a system architecture of the present application;
[0049] Figure 4 For the process of a cloud service;
[0050] Figure 5 Schematic diagram of the process of a text processing method provided by an embodiment of the present application;
[0051] Figure 6 Schematic diagram of a process for text sequence compression;
[0052] Figure 7 Schematic diagram of the process of a text processing method provided by an embodiment of the present application;
[0053] Figure 8 Schematic diagram of the beneficial effects of the present application;
[0054] Figure 9 Schematic diagram of the structure of a text processing device provided by an embodiment of the present application;
[0055] Figure 10 Schematic diagram of the structure of an execution device provided by an embodiment of the present application;
[0056] Figure 11 Schematic diagram of the structure of a training device provided by an embodiment of the present application;
[0057] Figure 12 Schematic diagram of the structure of a chip provided by an embodiment of the present application. Detailed implementation manners
[0058] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.
[0059] The embodiments of the present application will be described below with reference to the accompanying drawings. As can be known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0060] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0061] As used herein, the terms "substantially", "about" and similar terms are used as approximate terms, rather than terms of degree, and are intended to account for the inherent deviations of measured or calculated values that would be known to a person of ordinary skill in the art. In addition, the use of "may" when describing embodiments of the present application refers to "one or more possible embodiments". As used herein, the terms "use", "using", and "used" may be regarded as synonymous with the terms "utilize", "utilizing", and "utilized", respectively. Additionally, the term "exemplary" is intended to refer to an example or illustration.
[0062] First, the overall workflow of the artificial intelligence system will be described. Please refer to Figure 1A , Figure 1A Shown is a schematic structural diagram of an artificial intelligence main framework. The above artificial intelligence main framework will be described below from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0063] (1) Infrastructure
[0064] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through a basic platform. It communicates with the external world through sensors; the computing power is provided by intelligent chips (hardware acceleration chips such as CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external world to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.
[0065] (2) Data
[0066] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves Internet of Things data of traditional devices, including business data of existing systems and perception data such as force, displacement, liquid level, temperature, humidity, etc.
[0067] (3) Data Processing
[0068] Data processing generally includes data training, machine learning, deep learning, search, inference, decision-making, etc.
[0069] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.
[0070] Inference refers to the process of simulating the intelligent reasoning mode of humans in a computer or intelligent system, based on an inference control strategy, and using formalized information for machine thinking and problem-solving. The typical function is search and matching.
[0071] Decision-making refers to the process of making decisions after intelligent information is inferred, and usually provides functions such as classification, sorting, prediction, etc.
[0072] (4) General Capabilities
[0073] After the data undergoes the above-mentioned data processing, some general capabilities can be formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0074] (5) Intelligent Products and Industry Applications
[0075] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize intelligent information decision-making, and realize landing applications. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.
[0076] This application can be applied to the field of text processing in the field of artificial intelligence. Taking text processing as an example, multiple application scenarios implemented in products will be introduced below.
[0077] First, the application scenarios of this application will be introduced.
[0078] This application can be but is not limited to being applied to application programs with text generation functions (hereinafter can be simply referred to as text generation application programs) or cloud services provided by cloud-side servers, etc. Next, they will be introduced separately:
[0079] I. Text Generation Application Programs
[0080] The product form of the embodiments of this application can be a text generation application program. An application program with text generation function can run on a terminal device or a cloud-side server.
[0081] In a possible implementation, a text generation application can implement the task of text generation based on an input text generation request (including a text sequence to be processed), and obtain a processing result.
[0082] In particular, the text sequence to be processed is a very long text sequence.
[0083] Among them, the text generation request can indicate generative tasks such as summarization, dialogue, question and answer, translation, etc.
[0084] For example, the text generation request can be: Please summarize the following text into a summary within 100 words.
[0085] For example, the text generation request can be: Use the four characters "Pangu Zhizi" as the beginning of each sentence to create a seven-character quatrain poem with the rhyme of "a".
[0086] In a possible implementation, the user can open the text generation application installed on the terminal device and input a text processing instruction containing the text sequence to be processed. The text generation application can process the text processing instruction containing the text sequence to be processed through the method provided in the embodiments of the present application, and present the processing result to the user (the presentation method can be but is not limited to display, save, upload to the cloud side, etc.).
[0087] In a possible implementation, the user can open the text generation application installed on the terminal device and input a text processing instruction containing the text sequence to be processed. The text generation application can send the text processing instruction containing the text sequence to be processed to the server on the cloud side. The server on the cloud side processes the text processing instruction containing the text sequence to be processed through the method provided in the embodiments of the present application, and sends the processing result back to the terminal device. The terminal device can present the processing result to the user (the presentation method can be but is not limited to display, save, upload to the cloud side, etc.).
[0088] Optionally, the input and output can occur on a paid dialogue interface, a terminal application service that calls relevant interfaces, or a built-in function of the terminal device.
[0089] Next, the text generation application in the embodiments of the present application will be introduced respectively from the functional architecture and the product architecture for implementing the functions.
[0090] Refer to Figure 1B , Figure 1B which is a schematic diagram of the functional architecture of the text generation application in the embodiments of the present application:
[0091] In a possible implementation, as Figure 1BAs shown, the text generation application 102 can receive the input parameter 101 (e.g., a text processing instruction containing a text sequence to be processed) and generate a processing result 103. The text generation application 102 can be executed on (for example) at least one computer system and includes computer code that, when executed by one or more computers, causes the computers to execute the methods provided by the embodiments of the present application.
[0092] Referring to Figure 1C , Figure 1C This is a schematic diagram of the entity architecture for running the text generation application in the embodiments of the present application:
[0093] See Figure 1C , Figure 1C shows a schematic diagram of a system architecture. The system may include a terminal 100 and a server 200. Among them, the server 200 may include one or more servers ( Figure 1C illustrated by taking one server as an example), and the server 200 may provide the methods provided by the embodiments of the present application for one or more terminals.
[0094] Among them, a text generation application may be installed on the terminal 100. The above application and web page may provide an interface. The terminal 100 may receive relevant parameters input by the user on the text generation interface and send the above parameters to the server 200. The server 200 may obtain a processing result based on the received parameters and return the processing result to the terminal 100.
[0095] It should be understood that in some alternative implementations, the terminal 100 may also complete the action of obtaining the processing result based on the received parameters by itself without the cooperation of the server, which is not limited in the embodiments of the present application.
[0096] Next, describe Figure 1C the product form of the terminal 100 in
[0097] The terminal 100 in the embodiments of the present application may be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiments of the present application do not make any restrictions on this.
[0098] Figure 1D shows an optional schematic diagram of the hardware structure of the terminal 100.
[0099] Reference Figure 1D As shown, the terminal 100 may include components such as a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, a power supply 190, etc. Those skilled in the art can understand that Figure 1D This is merely an example of a terminal or a multifunctional device, and does not constitute a limitation on the terminal or the multifunctional device. It may include more or fewer components than shown in the figure, or combine certain components, or different components.
[0100] The input unit 130 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the portable multifunctional device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object such as a finger, a joint, a stylus, etc. on or near the touch screen), and drive corresponding connection devices according to a pre-set program. The touch screen can detect the touch actions of the user on the touch screen, convert the touch actions into touch signals and send them to the processor 170, and can receive and execute commands sent by the processor 170; the touch signals at least include contact coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch screen. In addition to the touch screen 131, the input unit 130 may further include other input devices. Specifically, the other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0101] Among them, the input device 132 can receive text processing instructions and the like including a text sequence to be processed input.
[0102] The display unit 140 can be used to display information input by the user or provided to the user, various menus of the terminal 100, an interactive interface, file display, and / or the playback of any multimedia file. In the embodiments of the present application, the display unit 140 can be used to display the interface of a text generation application, processing results, etc.
[0103] The memory 120 can be used to store instructions and data. The memory 120 mainly includes a storage instruction area and a storage data area. The storage data area can store various data, such as multimedia files, texts, etc.; the storage instruction area can store software units such as an operating system, applications, instructions required for at least one function, or subsets or extended sets thereof. It can also include a non-volatile random access memory; it provides the processor 170 with functions including managing the hardware, software, and data resources in the computing processing device, supporting control software and applications. It is also used for storing multimedia files, as well as storing running programs and applications.
[0104] The processor 170 is the control center of the terminal 100. It connects various parts of the entire terminal 100 through various interfaces and lines. By running or executing the instructions stored in the memory 120 and calling the data stored in the memory 120, it executes various functions of the terminal 100 and processes data, thereby controlling the terminal device as a whole. Optionally, the processor 170 may include one or more processing units; preferably, the processor 170 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 170 either. In some embodiments, the processor and the memory can be implemented on a single chip. In some embodiments, they can also be separately implemented on independent chips. The processor 170 can also be used to generate corresponding operation control signals, send them to corresponding components of the computing processing device, read and process data in software, especially read and process the data and programs in the memory 120, so that each function module therein executes corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.
[0105] Among them, the memory 120 can be used to store software codes related to the text processing method. The processor 170 can execute the steps of the text processing method of the chip, or can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to implement corresponding functions.
[0106] The radio frequency unit 110 (optional) can be used for receiving and transmitting information or signals during a call. For example, after receiving the downlink information from the base station, it is sent to the processor 170 for processing; in addition, the uplink data designed is sent to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 110 can also communicate with network devices and other devices through wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0107] Wherein, in the embodiment of the present application, the radio frequency unit 110 can send a text processing instruction including a text sequence to be processed to the server 200 and receive the processing result sent by the server 200.
[0108] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network interface.
[0109] The terminal 100 further includes a power supply 190 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 170 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.
[0110] The terminal 100 further includes an external interface 180, which can be a standard Micro USB interface or a multi-pin connector, and can be used to connect the terminal 100 to other devices for communication or to connect a charger to charge the terminal 100.
[0111] Although not shown, the terminal 100 may further include a flashlight, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be elaborated here. Some or all of the methods described below can be applied to the terminal 100 as Figure 1D shown.
[0112] Next, the product form of the server 200 will be described. Figure 1C in the server 200;
[0113] Figure 2 A schematic structural diagram of a server 200 is provided, as Figure 2 shown. The server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other through the bus 201.
[0114] The bus 201 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 2 only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0115] The processor 202 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0116] The memory 204 can include volatile memory, such as random access memory (RAM). The memory 204 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0117] Among them, the memory 204 can be used to store software codes related to the text processing method, and the processor 202 can execute the steps of the text processing method of the chip, or can also schedule other units to implement the corresponding functions.
[0118] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processing (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above-mentioned hardware system without the function of executing instructions and the hardware system with the function of executing instructions.
[0119] It should be understood that the steps related to the model inference process in the embodiments of the present application involve AI-related operations. When performing AI operations, the instruction execution architectures of the terminal device and the server are not limited to the architecture of the above-mentioned processor combined with the memory. The following will be combined with Figure 3 to introduce the system architecture provided by the embodiments of the present application in detail.
[0120] Figure 3 is a schematic diagram of the system architecture provided by the embodiments of the present application. As Figure 3 shown, the system architecture 500 includes an execution device 510, a training device 520, a database 530, a client device 540, a data storage system 550, and a data acquisition system 560.
[0121] The execution device 510 includes a computing module 511, an I / O interface 512, a preprocessing module 513, and a preprocessing module 514. The computing module 511 may include a target model / rule 501, and the preprocessing module 513 and the preprocessing module 514 are optional.
[0122] Among them, the execution device 510 can be the above-mentioned terminal device or server that runs text generation applications.
[0123] The data acquisition device 560 is used to acquire training samples. After acquiring the training samples, the data acquisition device 560 stores these training samples in the database 530.
[0124] The training device 520 can train the neural network to be trained (such as the language model in the embodiments of the present application, etc.) based on the training samples maintained in the database 530 to obtain the target model / rule 501.
[0125] It should be understood that the training device 520 can perform a pre-training process on the neural network to be trained based on the training samples maintained in the database 530, or perform fine-tuning of the model on the basis of pre-training.
[0126] It should be noted that in practical applications, the training samples maintained in the database 530 do not necessarily all come from the collection of the data acquisition device 560, and it is also possible to receive them from other devices. Additionally, it should be noted that the training device 520 does not necessarily train the target model / rule 501 entirely based on the training samples maintained in the database 530, and it is also possible to obtain training samples from the cloud or other places for model training. The above descriptions should not be construed as limitations on the embodiments of the present application.
[0127] The target model / rule 501 trained according to the training device 520 can be applied to different systems or devices, such as the Figure 3 execution device 510 shown, and the execution device 510 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, a vehicle-mounted terminal, etc., or it can also be a server, etc.
[0128] Specifically, the training device 520 can transfer the trained model to the execution device 510.
[0129] In Figure 3 , the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with external devices, and the user can input data (such as a text processing instruction including a text sequence to be processed, etc.) to the I / O interface 512 through the client device 540.
[0130] The preprocessing modules 513 and 514 are used to perform preprocessing on the input data received by the I / O interface 512. It should be understood that there may be no preprocessing modules 513 and 514 or only one preprocessing module. When the preprocessing modules 513 and 514 do not exist, the computing module 511 can directly process the input data.
[0131] During the preprocessing of the input data by the execution device 510, or during the relevant processing such as the computing module 511 of the execution device 510 performing calculations, the execution device 510 can call the data, code, etc. in the data storage system 550 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing in the data storage system 550.
[0132] Finally, the I / O interface 512 provides the processing result to the client device 540 and thus to the user.
[0133] In Figure 3 the illustrated case, the user can manually provide input data, and the "manually provided input data" can be operated through the interface provided by the I / O interface 512. In another case, the client device 540 can automatically send input data to the I / O interface 512. If the client device 540 is required to automatically send input data and user authorization is needed, the user can set corresponding permissions in the client device 540. The user can view the results output by the execution device 510 in the client device 540, and the specific presentation form can be display, sound, action, etc. The client device 540 can also be used as a data collection end to collect the input data input to the I / O interface 512 and the output result of the output I / O interface 512 as shown in the figure as new sample data and store it in the database 530. Of course, it is also possible not to collect through the client device 540, but for the I / O interface 512 to directly store the input data input to the I / O interface 512 and the output result of the output I / O interface 512 as shown in the figure as new sample data in the database 530.
[0134] It should be noted that Figure 3 this is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 3 it, the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the above execution device 510 can be deployed in the client device 540.
[0135] From the inference side of the model:
[0136] In an embodiment of the present application, the computing module 511 of the above execution device 520 can obtain the code stored in the data storage system 550 to implement the steps related to the sum model inference process in the embodiment of the present application.
[0137] In the embodiments of the present application, the computing module 511 of the execution device 520 may include hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above hardware systems without the function of executing instructions and hardware systems with the function of executing instructions.
[0138] Specifically, the computing module 511 of the execution device 520 may be a hardware system with the function of executing instructions. The steps related to the model inference process provided in the embodiments of the present application may be software codes stored in the memory. The computing module 511 of the execution device 520 may obtain the software codes from the memory and execute the obtained software codes to implement the steps related to the model inference process provided in the embodiments of the present application.
[0139] It should be understood that the computing module 511 of the execution device 520 may be a combination of a hardware system without the function of executing instructions and a hardware system with the function of executing instructions. Some of the steps related to the model inference process provided in the embodiments of the present application may also be implemented by the hardware system without the function of executing instructions in the computing module 511 of the execution device 520, which is not limited here.
[0140] From the perspective of the training side of the model:
[0141] In the embodiments of the present application, the above training device 520 may obtain the codes stored in the memory ( Figure 3 not shown in the figure, which may be integrated with the training device 520 or deployed separately from the training device 520) to implement the steps related to model training in the embodiments of the present application.
[0142] In the embodiments of the present application, the training device 520 may include a hardware circuit (such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with the function of executing instructions, such as a CPU, a DSP, etc., or a hardware system without the function of executing instructions, such as an ASIC, an FPGA, etc., or a combination of the above hardware systems without the function of executing instructions and hardware systems with the function of executing instructions.
[0143] It should be understood that the training device 520 may be a combination of a hardware system without the function of executing instructions and a hardware system with the function of executing instructions. Some steps related to the neutralization model training provided in the embodiments of the present application may also be implemented by the hardware system without the function of executing instructions in the training device 520, which is not limited herein.
[0144] II. Text generation cloud services provided by the server:
[0145] In a possible implementation, the server may provide a text generation service for the terminal side through an application programming interface (API).
[0146] Among them, the terminal device may send relevant parameters (such as a text processing instruction including a text sequence to be processed) to the server through the API provided by the cloud, and the server may obtain a processing result, etc. based on the received parameters, and return the processing result to the terminal.
[0147] The descriptions of the terminal and the server may refer to the descriptions of the above embodiments, which will not be elaborated here.
[0148] As Figure 4 shows the process of using a text generation cloud service provided by a cloud platform.
[0149] 1. Open and purchase the text generation service.
[0150] 2. Users can download the software development kit (SDK) of the text generation service. Usually, the cloud platform provides multiple development versions of the SDK for users to choose according to the requirements of the development environment, such as the SDK of JAVA version, the SDK of python version, the SDK of PHP version, the SDK of Android version, etc.
[0151] 3. After the user downloads the corresponding version of the SDK to the local according to the requirements, the SDK project is imported into the local development environment, and configured and debugged in the local development environment. Other functions can also be developed in the local development environment, so as to form an application integrating text generation capabilities.
[0152] 4. During the use of the text generation application, when text generation is required, the API call for text generation can be triggered. When the application triggers the text generation function, an API request is sent to the running instance of the text generation service in the cloud environment. Among them, the API request carries a text processing instruction containing the text sequence to be processed, and the running instance in the cloud environment processes the text processing instruction containing the text sequence to be processed to obtain the processing result.
[0153] 5. The cloud environment returns the processing result to the application, thus completing a method call provided by the embodiment of the present application.
[0154] Since the embodiments of the present application involve a large number of applications of neural networks, for the sake of easy understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application will be introduced below.
[0155] (1) Neural network
[0156] A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs (i.e., input data) and intercept 1 as inputs. The output of this operation unit can be:
[0157]
[0158] Among them, s = 1, 2, …… n, where n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neuron. f is the activation function of the neuron, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0159] (2) Deep Neural Network
[0160] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. Here, "many" does not have a specific measurement standard. Dividing DNN according to the positions of different layers, the neural network inside DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i + 1-th layer. Although DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply speaking, it is the following linear relationship expression: Among them, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer only performs such a simple operation on the input vector to obtain the output vector Since DNN has many layers, the number of coefficients W and the offset vector is also very large. The definitions of these parameters in DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts are the index 2 of the output third layer and the index 4 of the input second layer. In summary: The coefficient from the k-th neuron in the L - 1-th layer to the j-th neuron in the L-th layer is defined as Note that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can perform more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0161] (3) Backpropagation algorithm
[0162] The convolutional neural network can use the error backpropagation (BP) algorithm to correct the parameter sizes in the initial super-resolution model during training, making the reconstruction error loss of the super-resolution model smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial super-resolution model parameters are updated by backpropagating the error loss information, so that the error loss converges. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.
[0163] (4) Loss function
[0164] During the process of training a deep neural network, since we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the real target value or a value very close to the real target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0165] (6) Embedding: The low-dimensional embedding representation of high-dimensional words or sentences, so that words and sentences with similar meanings in the embedding vectors of words and sentences with close distances have similar meanings.
[0166] (7) Token: The elements generated after the text is tokenized.
[0167] (8)tokenizer: A tokenizer that tokenizes text according to a vocabulary and generates metrics.
[0168] (9)foundation model: A foundation model, usually referring to a large language model that has been pre-trained on a large scale with data from different domains.
[0169] (10)pre-training: Pre-training, which means learning from a large amount of (text) input information without annotation using a self-supervised method. The models used usually fall into autoregressive models and autoencoder models.
[0170] Large language models are neural network models with a large number of parameters trained on a large amount of text data. They can accurately understand the meaning of language text or generate specified natural language text, and effectively handle various natural language processing tasks such as translation, dialogue, and code generation. Therefore, they have great commercial value. However, current large language models often have limitations on the length of the input text, making it impossible to effectively deploy them in long text scenarios such as books, academic papers, and legal contracts.
[0171] There are usually two main reasons for the limitation on the length of the input text of language models. One is that language models are often trained on short texts, and the positional encoding of large language models does not directly adapt to the situation of long texts, resulting in a significant decline in the performance of large language models on long texts. The other is that the inference cost of language models grows quadratically with the length of the input text, and the longer the text, the higher the load on the inference engine and platform.
[0172] Therefore, there is an urgent need for a method to increase the effective input length of language models.
[0173] To solve the above problems, the embodiments of the present application provide a text processing method. The text processing method of the embodiments of the present application will be introduced in detail below with reference to the accompanying drawings.
[0174] Refer to Figure 5 , Figure 5 which is a schematic flow of a text processing method provided by the embodiments of the present application. As Figure 5 shown, the text processing method provided by the embodiments of the present application can be, but is not limited to, the feed-forward process during model pre-training, the feed-forward process during model fine-tuning, the model inference process, etc. Specifically, the method may include steps 501 to 502, which will be described in detail below.
[0175] 501. Obtain a text sequence; the text sequence includes a plurality of sub-text sequences.
[0176] During the feed-forward process of model training (such as pre-training or fine-tuning), the text sequence can be the training samples that the language model needs to process.
[0177] During the feed-forward process of model training, the text sequence can be the training samples that the language model needs to process.
[0178] During the process of model inference, the text sequence can be the text sequence to be processed by the language model. For example, the text sequence can be carried in a text processing request.
[0179] In a possible implementation, the length of the text sequence is relatively long and exceeds the processing length supported by the input of the language model. Therefore, it is necessary to compress the text sequence.
[0180] In the embodiments of the present application, the text sequence can be divided (or can be called segmented) to obtain multiple sub-text sequences. For example, the text sequence can be a text paragraph including multiple sentences, and the text sequence can be segmented at the sentence level, that is, each sub-text sequence can include one or more sentences. For example, it can be controlled that the number of characters in each sub-text sequence obtained after segmentation is less than or equal to a preset value.
[0181] Among them, there can be overlap or complete stagger between different sub-text sequences after division. The existence of overlap can improve the continuity of the segmented content. For example, there can be overlap between at least two of the multiple sub-text sequences obtained after dividing the text sequence.
[0182] In a possible implementation, the text sequence includes m sentences, and the sub-text sequence includes n sentences, where n < m.
[0183] Exemplarily, the text sequence can include 10 sentences:
[0184] Sentence 1, Sentence 2, Sentence 3, Sentence 4, Sentence 5, Sentence 6, Sentence 7, Sentence 8, Sentence 9, Sentence 10.
[0185] When the sub-text sequences are completely staggered, the multiple sub-text sequences obtained after division can include the following sub-text sequences:
[0186] Sub-text sequence 1: Sentence 1, Sentence 2, Sentence 3;
[0187] Sub-text sequence 2: Sentence 4, Sentence 5, Sentence 6, Sentence 7;
[0188] Sub-text sequence 3: Sentence 8, Sentence 9.
[0189] When there is overlap between the sub-text sequences, the multiple sub-text sequences obtained after division can include the following sub-text sequences:
[0190] Sub - text sequence 1: Sentence 1, Sentence 2, Sentence 3, Sentence 4;
[0191] Sub - text sequence 2: Sentence 4, Sentence 5, Sentence 6, Sentence 7;
[0192] Sub - text sequence 3: Sentence 8, Sentence 9.
[0193] Among them, there is an overlapping part (Sentence 4) between Sub - text sequence 1 and Sub - text sequence 2.
[0194] 502. Group the multiple sub - text sequences to obtain multiple text groups; the semantic similarity between the sub - text sequences included in the same text group is greater than the first threshold; the multiple text groups are used to obtain model outputs through a language model.
[0195] The embodiments of this application hope to extend the length limit of the input text that the language model can process and improve the effective input length of the language model. There are usually two main reasons for the length limit. One is that language models are often trained on short texts, and the positional encoding of large language models does not directly adapt to the situation of long texts, resulting in a significant decline in the performance of large language models on long texts. The other is that the inference cost of the language model grows quadratically with the length of the input text, and the longer the text, the higher the load on the inference engine and the platform.
[0196] Therefore, it is necessary to perform semantic compression on the input text. However, on the one hand, there is also a problem of input text length limit when performing model compression. Therefore, before performing semantic compression, the text sequence can be divided, and each sub - text sequence obtained after division is compressed separately, and then the processing results are spliced, which is equivalent to compressing in a way of divide - and - conquer and then merging.
[0197] However, the division of the text sequence will destroy the coherence in the text, resulting in the loss of a lot of text semantic information when compressing each sub - text sequence. Therefore, the idea of this application is to design a grouping method for sub - text sequences, which can make the sub - text sequences included in each text group describe similar semantic content, maintain the coherence of the content of each sub - text sequence, thereby reducing the loss of text semantic information during compression, improving the compression effect, and further improving the processing effect of the language model.
[0198] Since the text sequences of natural languages often have a hierarchical structure. For example, books often have a chapter structure, where each chapter tells a complete content, and the chapters together form a complete book. Although not all texts have a chapter structure, the chapter structure can be implicitly modeled with topic paragraphs. Each topic paragraph describes similar content, which is different from the content before and after. If a long text is divided according to topic paragraphs, then the content of each topic can be compressed independently while maintaining the coherence of the content. After summarizing the compressed content in order, a better semantic processing result of the long text can be obtained.
[0199] Each of the above-mentioned divided topic paragraphs can be a text group in the embodiments of the present application. That is to say, the sub-text sequences in a text group can express the same theme.
[0200] In a possible implementation, the semantic similarity between the sub-text sequences included in the same text group is greater than a first threshold, while the semantic similarity between the sub-text sequences included in different text groups is less than the first threshold. That is to say, each text group obtained after division can include sub-text sequences with similar semantics, and the semantic differences between the sub-text sequences included in different text groups are relatively large.
[0201] Next, an example of a method for grouping multiple sub-text sequences is introduced:
[0202] In a possible implementation, a pre-trained embedding model can be used to perform embedding processing on each sub-text sequence to obtain embedding vectors. A weighted graph can be constructed based on the embedding vectors of multiple sub-text sequences. The similarity between the embedding vectors is the connection strength between the segments in the graph (this connection strength is represented by a weight value, so this graph structure can be called a weighted graph). Based on this, the modeling of the weighted graph between sub-text sequences can be obtained. After obtaining the weighted graph of the long text, the continuous community structure in the graph can be found, which is equivalent to finding the division of the main paragraphs. Each community structure corresponds to a text group. The internal connections of the community structure are close, while the external connections are weak, which is consistent with the idea of implicitly modeled main paragraphs. That is to say, each community structure can include multiple graph nodes, each graph node corresponds to a sub-text sequence, and the similarity between the sub-text sequences within each community structure is relatively high, while the similarity between the sub-text sequences in different community structures is relatively low.
[0203] In addition, to ensure a relatively high continuity among the sub-text sequences within the same text group (i.e., their positions in the original text sequence are consecutive or very close, which allows the compression process to recognize richer continuous semantic information and thus improve the compression quality), when grouping the sub-text sequences, in addition to considering semantic similarity, the position similarity of the sub-text sequences in the text sequence can also be taken into account. That is, through grouping, the position similarity of the sub-text sequences included in the same text group in the text sequence is greater than a second threshold. In addition, it is also possible to group in such a way that the position similarity of the sub-text sequences included in the same text group in the text sequence is less than the second threshold.
[0204] Next, an example of a way to group multiple sub-text sequences is introduced:
[0205] Taking the implementation method of the above weighted graph as an example, in order to obtain a continuous community structure, the connections between nodes that are far apart can be masked. For example, the weight value of the connection between nodes that are far apart in the graph can be set to 0. This makes the nodes included in the main paragraph also adjacent and arranged in order for segmentation.
[0206] Among them, the multiple text groups are used to obtain a model output through a language model.
[0207] Among them, after obtaining the multiple text groups, other processes such as compression and sequence expansion can also be performed on each text group.
[0208] Taking compression as an example, the embodiments of the present application can compress each of the multiple text groups to obtain multiple processing results. Furthermore, the multiple text groups can be used to obtain a model output through a language model, including: the multiple processing results are used to obtain a model output through a language model, and the multiple processing results are used to obtain a processing result of the text sequence through a language model.
[0209] In a possible implementation, the compression is semantic compression for reducing the text length. By using semantic compression, while retaining important semantic information, the long text input to the language model is compressed into a shorter natural language text, which equivalently extends the input length of the large language model. In addition, the sub-text sequences within each text group obtained after grouping in the embodiments of the present application describe similar semantic content, which can maintain the coherence of the content of each sub-text sequence, thereby reducing the loss of text semantic information during compression and improving the compression effect.
[0210] For example, according to the multiple text groups, multiple processing results can be obtained through the language model or other machine learning models other than the language model. Machine learning for compression also has limitations on the length of the input text. Through the method of this application, the text sequence is grouped, and the long text input to the compression model during compression is changed into shorter text groups, which equivalently extends the input length of the compression model. Moreover, each text group includes sub-text sequences describing similar semantic content, maintaining the coherence of the content of each sub-text sequence, thereby reducing the loss of text semantic information during compression and improving the compression effect.
[0211] According to the positional relationship of the multiple text groups among the text sequences, the multiple processing results are spliced to obtain the spliced processing result. After summarizing the compressed content in order, a semantic processing result of a long text with better effect can be obtained.
[0212] In a possible implementation, according to the multiple processing results, the processing result of the text sequence can be obtained through a language model.
[0213] Optionally, the multiple processing results can be extended through a sequence extension method, and then according to the extended multiple processing results, the processing result of the text sequence can be obtained through a language model.
[0214] Among them, the language model can be a large language model (LLM). The language model can perform text processing tasks to obtain processing results. For example, the text processing tasks can include but are not limited to abstract generation, question answering, text translation, etc.
[0215] Next, a specific example is used to introduce a process of an embodiment of this application.
[0216] Refer to Figure 6, first, the long text is divided into sentence-level segments. Here, the segmentation can control the number of characters for segmentation, or overlap the characters between segments to ensure the continuity of the segmented content. After obtaining the sequential segmentation, a pre-trained sentence embedding model can be used to embed the segments into vectors. The similarity between the vectors is the connection strength between the segments in the graph, and thus the weighted graph modeling of the text is obtained. After obtaining the graph structure of the long text, finding the continuous community structure in the graph is equivalent to finding the topic paragraph division. The internal connections within the community structure are close, while the connections with the outside are weak, which is consistent with the idea of implicitly modeled topic paragraphs. To obtain a continuous community structure, the connections between segments that are far apart are masked, that is, their values in the graph are set to 0. In this way, by using the segmentation algorithm on the graph, the segmented blocks obtained are the required topic paragraphs, and thus the topic paragraphs are obtained, and the topic paragraphs also contain adjacent segments arranged in sequence. After model compression and sequential summarization, the semantic processing result of the long text is obtained. Although the length is reduced, the key semantic information is retained.
[0217] The embodiments of the present application can be applicable to the scenario where the large language model processes texts exceeding its maximum text length, and can also help reduce the scenario where the large model inference encounters video memory limitations. Figure 7 It is a schematic diagram of an application architecture of the embodiments of the present application.
[0218] Among them, the tokenizer can be used to specify the action of dividing the text sequence into multiple sub-text sequences. The main division can perform the action of grouping multiple sub-text sequences. The chunk summary can perform the action of compressing each text group. The text recombination can perform the action of splicing multiple processing results.
[0219] The embodiments of the present application can effectively extend the input length of the large language model in natural language tasks. The following is the gain result compared with the dataset and other extension methods.
[0220] Table 1
[0221]
[0222] The long text summarization task can be completed by directly chunking and summarizing according to the input length of the large language model and then summarizing. And the technical solution herein is a natural extension of this solution from the semantic perspective. As shown in the results in Table 1, the solution provided by the embodiments of the present application has obvious gain effects compared with the existing solutions.
[0223] The Q&A task is more complex than the summarization task and cannot be processed by splitting and then merging. All the information must be processed at once. The key retrieval task is a Q&A task that can be generated to any length and is a commonly used dataset for testing the input length extension scheme of large models. Referring to Table 2, the results of Llama 2 in Table 2 show that Llama 2 has an input limit of 4,000 tokens, and the performance of the model drops sharply after exceeding the limit. After applying this solution to Llama 2, the input length can be extended to 30,000 tokens, and the accuracy is over 90%. This solution can also easily stack other extension schemes to further increase the input text length limit.
[0224] Table 2
[0225]
[0226] LongBench is a comprehensive long text task dataset, including four types of tasks: single text Q&A, multi-document Q&A, summarization, and few-shot learning. We evaluated three English datasets for each type. Figure 8 The results in show that the method provided by the embodiments of the present application exceeds the optimal academic extension method Yarn in most tasks and exceeds the original model in most tasks.
[0227] Referring to Figure 9 , Figure 9 is a structural schematic of a text processing device provided by an embodiment of the present application. As Figure 9 shown, a text processing device 900 provided by an embodiment of the present application includes:
[0228] An acquisition module 901, configured to acquire a text sequence; the text sequence includes a plurality of sub-text sequences;
[0229] Specific descriptions of the acquisition module 901 can refer to the description of step 501 in the above embodiments and will not be elaborated here.
[0230] A grouping module 902, configured to group the plurality of sub-text sequences to obtain a plurality of text groups; the semantic similarity between the sub-text sequences included in the same text group is greater than a first threshold; the plurality of text groups are used to obtain a model output through a language model. Specific descriptions of the grouping module 902 can refer to the description of step 502 in the above embodiments and will not be elaborated here.
[0231] In a possible implementation, the device further includes:
[0232] A compression module 903, configured to compress each text group in the plurality of text groups to obtain a plurality of processing results;
[0233] The multiple text groups are used to obtain a model output through a language model, including:
[0234] The multiple processing results are used to obtain a model output through a language model.
[0235] In a possible implementation, the semantic similarity between the sub-text sequences included in different text groups is less than the first threshold.
[0236] In a possible implementation, the positional similarity between the sub-text sequences included in the same text group in the text sequence is greater than a second threshold.
[0237] In a possible implementation, the positional similarity between the sub-text sequences included in the same text group in the text sequence is less than the second threshold.
[0238] In a possible implementation, there is an overlap between at least two of the multiple sub-text sequences.
[0239] In a possible implementation, the compression module 903 is further configured to:
[0240] Splice the multiple processing results according to the positional relationship between the multiple text groups in the text sequence to obtain a spliced processing result.
[0241] In a possible implementation, the device further includes:
[0242] A task processing module, configured to obtain a processing result of the text sequence through a language model according to the multiple text groups.
[0243] In a possible implementation, the compression module 903 is specifically configured to:
[0244] Compress each of the multiple text groups through the language model or other machine learning models other than the language model to obtain multiple processing results.
[0245] In a possible implementation, the model output is the result of performing one of the following tasks on the text sequence:
[0246] Abstract extraction task, text generation task, dialogue task, question and answer task, text translation task, knowledge retrieval task.
[0247] Next, an execution device provided in an embodiment of the present application will be introduced. Please refer to Figure 10 , Figure 10A schematic structural diagram of an execution device provided by an embodiment of the present application. The execution device 1000 may specifically be embodied as a virtual reality (VR) device, a mobile phone, a tablet computer, a laptop computer, a smart wearable device, a monitoring data processing device, a server, etc., which is not limited herein. Specifically, the execution device 1000 includes: a receiver 1001, a transmitter 1002, a processor 1003, and a memory 1004 (where the number of processors 1003 in the execution device 1000 may be one or more, Figure 10 and one processor is taken as an example herein). Among them, the processor 1003 may include an application processor 10031 and a communication processor 10032. In some embodiments of the present application, the receiver 1001, the transmitter 1002, the processor 1003, and the memory 1004 may be connected through a bus or other means.
[0248] The memory 1004 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1003. A part of the memory 1004 may also include a non-volatile random access memory (NVRAM). The memory 1004 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions may include various operation instructions for implementing various operations.
[0249] The processor 1003 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, where the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to a data bus. However, for the sake of clear illustration, all kinds of buses are referred to as a bus system in the figure.
[0250] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1003. The processor 1003 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1003 or instructions in the form of software. The above-mentioned processor 1003 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1003 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1004, and the processor 1003 reads the information in the memory 1004 and combines its hardware to complete the steps involved in the model inference process in the above method.
[0251] The receiver 1001 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1002 can be used to output digital or character information through the first interface; the transmitter 1002 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1302 can also include a display device such as a display screen.
[0252] The embodiments of the present application also provide a training device. Please refer to Figure 11 , Figure 11It is a schematic structural diagram of a training device provided by an embodiment of the present application. Specifically, the training device 1100 is implemented by one or more servers. The training device 1100 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1111 (for example, one or more processors) and a memory 1132, and one or more storage media 1130 (for example, one or more mass storage devices) for storing application programs 1142 or data 1144. Among them, the memory 1132 and the storage medium 1130 may be transient storage or persistent storage. The program stored in the storage medium 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the training device. Further, the central processing unit 1111 may be configured to communicate with the storage medium 1130 and execute a series of instruction operations in the storage medium 1130 on the training device 1100.
[0253] The training device 1100 may further include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, and one or more input / output interfaces 1158; or, one or more operating systems 1141, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0254] In an embodiment of the present application, the central processing unit 1111 is used to execute the actions related to model training in the above embodiments.
[0255] An embodiment of the present application also provides a computer program product, which when running on a computer, causes the computer to execute the steps executed by the aforementioned execution device, or causes the computer to execute the steps executed by the aforementioned training device.
[0256] An embodiment of the present application also provides a computer-readable storage medium, in which a program for signal processing is stored, and when it runs on a computer, it causes the computer to execute the steps executed by the aforementioned execution device, or causes the computer to execute the steps executed by the aforementioned training device.
[0257] The execution device, training device, or terminal device provided by the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, or the like. The processing unit may execute the computer-executable instructions stored in the storage unit to cause the chip in the execution device to execute the text processing method described in the above embodiments, or to cause the chip in the training device to execute the text processing method described in the above embodiments. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0258] Specifically, please refer to Figure 12 , Figure 12 which is a schematic structural diagram of the chip provided by the embodiments of the present application. The chip may be embodied as a neural network processor NPU 1200, and the NPU 1200 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 1203, and the arithmetic circuit 1203 is controlled by the controller 1204 to extract matrix data from the memory and perform multiplication operations.
[0259] In some implementations, the arithmetic circuit 1203 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 1203 is a two-dimensional systolic array. The arithmetic circuit 1203 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1203 is a general matrix processor.
[0260] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1202 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1201 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are saved in the accumulator 1208.
[0261] The unified memory 1206 is used to store input data and output data. The weight data directly passes through the Direct Memory Access Controller (DMAC) 1205 and is transferred to the weight memory 1202 by the DMAC. The input data is also transferred to the unified memory 1206 by the DMAC.
[0262] The BIU is the Bus Interface Unit, that is, the bus interface unit 1210, which is used for the interaction between the AXI bus, the DMAC, and the Instruction Fetch Buffer (IFB) 1209.
[0263] The bus interface unit 1210 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch buffer 1209 to obtain instructions from the external memory, and is also used for the storage unit access controller 1205 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0264] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1206, or transfer the weight data to the weight memory 1202, or transfer the input data to the input memory 1201.
[0265] The vector calculation unit 1207 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit 1203, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.
[0266] In some implementations, the vector calculation unit 1207 can store the processed output vector in the unified memory 1206. For example, the vector calculation unit 1207 can apply a linear function; or, a non-linear function to the output of the arithmetic circuit 1203, such as performing linear interpolation on the feature plane extracted by the convolutional layer, or, for example, a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 1207 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 1203, such as for use in subsequent layers in a neural network.
[0267] The instruction fetch buffer 1209 connected to the controller 1204 is used to store the instructions used by the controller 1204;
[0268] The unified memory 1206, the input memory 1201, the weight memory 1202, and the fetch memory 1209 are all On-Chip memories. The external memory is private to the NPU hardware architecture.
[0269] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above programs.
[0270] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0271] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by means of dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of this application.
[0272] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0273] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. A text processing method, characterized in that, The method includes: Obtaining a text sequence; the text sequence includes a plurality of sub - text sequences; Grouping the plurality of sub - text sequences to obtain a plurality of text groups; the semantic similarity between the sub - text sequences included in the same text group is greater than a first threshold; the plurality of text groups are used to obtain a model output through a language model.
2. The method according to claim 1, characterized in that, The method further includes: Compressing each of the plurality of text groups to obtain a plurality of processing results; The plurality of text groups are used to obtain a model output through a language model, including: The plurality of processing results are used to obtain a model output through a language model.
3. The method according to claim 1 or 2, characterized in that, The semantic similarity between the sub - text sequences included in different text groups is less than the first threshold.
4. The method according to any one of claims 1 to 3, characterized in that, The positional similarity between the sub - text sequences included in the same text group in the text sequence is greater than a second threshold.
5. The method according to any one of claims 1 to 4, characterized in that, The positional similarity between the sub - text sequences included in the same text group in the text sequence is less than the second threshold.
6. The method according to any one of claims 1 to 5, characterized in that, There is an overlap between at least two of the plurality of sub - text sequences.
7. The method according to any one of claims 2 to 6, characterized in that, The method further includes: Splicing the plurality of processing results according to the positional relationship between the plurality of text groups in the text sequence to obtain a spliced processing result.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Obtaining a processing result of the text sequence through a language model according to the plurality of text groups.
9. The method according to any one of claims 2 to 8, characterized in that, The compressing each of the plurality of text groups to obtain a plurality of processing results includes: Compressing each of the plurality of text groups through the language model or other machine learning models other than the language model to obtain a plurality of processing results.
10. The method according to any one of claims 1 to 9, characterized in that, The model output is the result of performing one of the following tasks on the text sequence: Abstract extraction task, text generation task, dialogue task, question - answering task, text translation task, knowledge retrieval task.
11. A text processing device, characterized in that, The apparatus includes: An obtaining module, configured to obtain a text sequence; the text sequence includes a plurality of sub - text sequences; A grouping module, configured to group the plurality of sub - text sequences to obtain a plurality of text groups; the semantic similarity between the sub - text sequences included in the same text group is greater than a first threshold; the plurality of text groups are used to obtain a model output through a language model.
12. The device according to claim 11, characterized in that, The apparatus further includes: A compression module, configured to compress each of the plurality of text groups to obtain a plurality of processing results; The plurality of text groups are used to obtain a model output through a language model, including: The plurality of processing results are used to obtain a model output through a language model.
13. The device according to claim 11 or 12, characterized in that, The semantic similarity between the sub - text sequences included in different text groups is less than the first threshold.
14. The device according to any one of claims 11 to 13, characterized in that The positional similarity between the sub - text sequences included in the same text group in the text sequence is greater than a second threshold.
15. The device according to any one of claims 11 to 14, characterized in that The positional similarity between the sub - text sequences included in the same text group in the text sequence is less than the second threshold.
16. The device according to any one of claims 11 to 15, characterized in that There is an overlap between at least two of the plurality of sub - text sequences.
17. The device according to any one of claims 12 to 16, characterized in that The compression module is further configured to: Splice the plurality of processing results according to the positional relationship between the plurality of text groups in the text sequence to obtain a spliced processing result.
18. The device according to any one of claims 11 to 17, characterized in that The apparatus further includes: A task processing module, configured to obtain a processing result of the text sequence according to the multiple text groups through a language model.
19. The device according to any one of claims 11 to 18, characterized in that The model output is a result of performing one of the following tasks on the text sequence: Abstract extraction task, text generation task, dialogue task, question and answer task, text translation task, knowledge retrieval task.
20. The device according to any one of claims 12 to 19, characterized in that The compression module is specifically configured to: Compress each of the multiple text groups through the language model or other machine learning models other than the language model to obtain multiple processing results.
21. A computer storage medium, characterized in that The computer storage medium stores one or more instructions, and when the instructions are executed by one or more computers, the one or more computers perform the operations of the method according to any one of claims 1 to 10.
22. A computer program product, characterized in that Including computer-readable instructions, when the computer-readable instructions run on a computer device, the computer device executes the method according to any one of claims 1 to 10.
23. A system, comprising at least one processor and at least one memory; the processor and the memory are connected through a communication bus and communicate with each other; The at least one memory is used for storing code; The at least one processor is used for executing the code to execute the method according to any one of claims 1 to 10.
24. A chip, comprising a processor, characterized in that The processor is used to support the text processing device to implement the method according to any one of claims 1 to 10.
Citation Information
Cited By
Text processing method and apparatus using same
EP4814780A1
Text processing method and apparatus using same
WO2025124301A1