Data processing method and apparatus therefor

By inserting non-space string separators in the natural language processing system and deleting them during post-processing, the problems of word segmentation inconsistency and lossy word segmentation are solved, and the effect of lossless word segmentation is achieved.

WO2025130968A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140542
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-19
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing natural language processing system ignores the inconsistency caused by language characteristics in the word segmentation process, and often introduces noise to cause lossy word segmentation.

Method used

Lossless word segmentation is achieved by inserting non-space string separators into text, identifying the position of text separation, and accurately deleting the inserted separators during post-processing.

Benefits of technology

The problems of word segmentation inconsistency and lossy word segmentation are solved, and the original text is accurately restored during post-processing is achieved, which improves the accuracy and consistency of word segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140542_26062025_PF_FP_ABST
    Figure CN2024140542_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing method, relating to the field of artificial intelligence, and comprising: acquiring a text; inserting a separator into the text, the separator being a non-space character string, and the separator identifying a position of character separation in the text; after the separator has been inserted, performing word segmentation on the text, to obtain a word segmentation result. The separator in the present application is a non-space character string (or may be referred to as a character string that is not entirely spaces). Since the non-space characters are not commonly seen in natural language, the inserted separator can be accurately deleted during post-processing, thereby realizing lossless word segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

A data processing method and device thereof

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 22, 2023, with application number 202311786847.2 and application name “A data processing method and device thereof”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and in particular to a data processing method and device thereof. Background Art

[0003] Existing natural language processing systems use Byte Pair Encoding (BPE) or SentencePiece algorithms to process input text. This segmentation technique uses spaces as delimiters to significantly reduce the occurrence of unrecognized words during inference and improve the generalization and error tolerance of the language model. Without segmentation, all words in the language must appear in the training text. Otherwise, missing words (such as misspellings) will be treated as unrecognized words during inference, significantly impacting the language model's error tolerance and inference capabilities. Furthermore, without word segmentation, some words with the same origin but different word forms will be treated as independent by the language model. This prevents the shared semantics of these words from being fully learned, reducing the language model's generalization.

[0004] However, these algorithms focus more on building mathematical models to achieve a more optimal subword distribution, while (partially) ignoring the inconsistencies in word segmentation caused by the inherent characteristics of language. This means that the same word may have different segmentation results in different scenarios. A common solution is to introduce a text pre-processing (or post-processing) module before or after word segmentation. While this approach can effectively address the inconsistency issue, it often introduces noise, leading to lossy word segmentation. This means that the input text cannot be restored after pre-processing, word segmentation, reverse segmentation, and post-processing. Summary of the Invention

[0005] The present application provides a data processing method that can accurately delete inserted delimiters during post-processing, thereby achieving lossless word segmentation.

[0006] In a first aspect, the present application provides a data processing method, which includes: obtaining text; inserting a delimiter into the text, wherein the delimiter is a non-space string, and the delimiter is used to identify the position of the character separation in the text; and performing word segmentation on the text after the delimiter is inserted to obtain a word segmentation result.

[0007] In an embodiment of the present application, a non-space character string (or a character string that is not entirely space) is defined as a delimiter. For example, a character string that starts with a single or multiple non-space characters and ends with a single or multiple space characters is used as a delimiter. Since non-space characters are not as common as spaces in natural language, the inserted delimiter can be accurately deleted during post-processing, thereby achieving lossless word segmentation.

[0008] In a possible implementation, the delimiter is a character string that starts with a non-space character string and ends with at least one space.

[0009] In a possible implementation, the separator is a character corresponding to the source code "\u2582".

[0010] In a possible implementation, inserting a separator into the text includes:

[0011] Inserting a delimiter into the text based on at least one of the following methods:

[0012] inserting a separator at an adjacent position after a punctuation mark included in the text;

[0013] Inserting separators at locations in the text where different languages ​​are converted and not separated by spaces;

[0014] A delimiter is inserted at a position immediately after the last of the consecutive spaces included in the text.

[0015] In a possible implementation, the method further includes: obtaining a first processing result of the text through a language model according to the word segmentation result.

[0016] In addition, the word segmentation results can be directly post-processed (for example, including reverse word segmentation, separator removal, and character protection structure removal).

[0017] In a possible implementation, the first processing result includes a plurality of word segmentation units and the separator; and the method further includes:

[0018] Reverse word segmentation is performed on the first processing result, and the separator is removed to obtain a second processing result of the text.

[0019] In a possible implementation, a string of characters that are the same as the delimiter may also exist in the text. However, this part of the string belongs to the text and is not used as a delimiter. During post-processing, this part of the characters should not be deleted. Therefore, the string of characters that are the same as the delimiter can be protected in the text, that is, a string embedding structure is constructed to protect the non-space character portion of the above delimiter that appears in the input text, forming a protected string. Furthermore, the text after the delimiter is inserted also includes: a protected string, which is used to identify the text in the text that is the same as the delimiter, and indicates that the text that is the same as the delimiter is not a delimiter.

[0020] In a possible implementation, the first processing result further includes the protected character string; and the method further includes: removing the protected character string from the first processing result.

[0021] In a second aspect, the present application provides a data processing device, comprising:

[0022] Acquisition module, used to obtain text;

[0023] The processing module is used to insert a separator into the text, where the separator is a non-space character string and is used to identify the position where the characters in the text are separated; and perform word segmentation on the text after the separator is inserted to obtain a word segmentation result.

[0024] In a possible implementation, the separator is a character corresponding to the source code "\u2582".

[0025] In a possible implementation, the processing module is specifically configured to:

[0026] Inserting a separator into the text based on at least one of the following means:

[0027] inserting a separator at an adjacent position after a punctuation mark included in the text;

[0028] Inserting separators at locations in the text where different languages ​​are converted and not separated by spaces;

[0029] A delimiter is inserted at a position immediately after the last of the consecutive spaces included in the text.

[0030] In a possible implementation, the processing module is further configured to:

[0031] According to the word segmentation result, a first processing result of the text is obtained through a language model.

[0032] In a possible implementation, the first processing result includes a plurality of word segmentation units and the separator; and the processing module is further configured to:

[0033] Reverse word segmentation is performed on the first processing result, and the separator is removed to obtain a second processing result of the text.

[0034] In a possible implementation, the text after the separator is inserted further includes: a protection string, wherein the protection string is used to identify text in the text that has the same characters as the separator and indicates that the text that has the same characters as the separator is not a separator.

[0035] In a possible implementation, the first processing result also includes the protected character string; and the processing module is further configured to:

[0036] The protected character string is removed from the first processing result.

[0037] In a third aspect, an embodiment of the present application provides a data processing device, which may include a memory, a processor, and a bus system, wherein the memory is used to store programs, and the processor is used to execute the programs in the memory to perform the first aspect and any optional method thereof.

[0038] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned first aspect and any optional method thereof.

[0039] In a fifth aspect, an embodiment of the present application provides a computer program, which, when executed on a computer, enables the computer to execute the above-mentioned first aspect and any optional method thereof.

[0040] In a sixth aspect, the present application provides a chip system comprising a processor configured to support a data processing device in implementing the functions described in the aforementioned aspects, such as transmitting or processing the data or information described in the aforementioned methods. In one possible design, the chip system further comprises a memory configured to store program instructions and data necessary for executing or training the device. The chip system may consist solely of a chip or may include a chip and other discrete components. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] FIG1A is a schematic diagram of a structure of an artificial intelligence main framework;

[0042] 1B and 1C are schematic diagrams of the application system framework of the present invention;

[0043] FIG1D is a schematic diagram of an optional hardware structure of a terminal;

[0044] FIG2 is a schematic diagram of the structure of a server;

[0045] Figures 3 to 5 are schematic diagrams of a system architecture of the present application;

[0046] Figure 6 shows a cloud service process;

[0047] Figure 7 is a process of cloud services;

[0048] FIG8 is a flowchart of a data processing method provided in an embodiment of the present application;

[0049] FIG9 is a flowchart of a data processing method provided in an embodiment of the present application;

[0050] FIG10 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;

[0051] FIG11 is a schematic diagram of the structure of an execution device provided in an embodiment of the present application;

[0052] FIG12 is a schematic diagram of a structure of a training device provided in an embodiment of the present application;

[0053] FIG13 is a schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] The following describes the embodiments of the present invention in conjunction with the accompanying drawings. The terms used in the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.

[0055] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0056] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0057] As used herein, the terms "substantially," "about," and similar terms are used as terms of approximation, not as terms of degree, and are intended to take into account the inherent variations in measurements or calculations that one of ordinary skill in the art would recognize. Furthermore, the use of "may" when describing embodiments of the present invention refers to "one or more possible embodiments." As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively. Additionally, the term "exemplary" is intended to refer to an example or illustration.

[0058] First, let's describe the overall workflow of an AI system. See Figure 1A, which shows a schematic diagram of the main AI framework. This AI framework will be explained from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.

[0059] (1) Infrastructure

[0060] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.

[0061] (2) Data

[0062] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0063] (3) Data processing

[0064] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0065] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.

[0066] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.

[0067] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.

[0068] (4) General ability

[0069] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0070] (5) Smart products and industry applications

[0071] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.

[0072] This application can be applied to the field of natural language processing in the field of artificial intelligence. Taking natural language processing as an example, the following will introduce multiple application scenarios that have been implemented in products.

[0073] First, we will introduce the application scenarios of this application. This application can be applied to, but is not limited to, applications with natural language processing functions (hereinafter referred to as natural language processing applications) or cloud services provided by cloud-side servers. The following are introduced separately:

[0074] 1. Natural Language Processing Applications

[0075] The product form of the embodiment of the present application can be a natural language processing application. The natural language processing application can be run on a terminal device or a cloud-side server.

[0076] In one possible implementation, a natural language processing application can implement natural language processing tasks, such as summary generation, text reply, text generation, and the like.

[0077] It should be understood that natural language processing can be implemented based on a language model, and the present application can be more specifically applied to the word segmentation process before inputting the language model.

[0078] In one possible implementation, a user can open a natural language processing application installed on a terminal device and enter text. The natural language processing application can obtain the natural language processing results through the method provided in the embodiment of the present application and present the natural language processing results to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).

[0079] In one possible implementation, a user can open a natural language processing application installed on a terminal device and enter text. The natural language processing application can perform natural language processing on the text using the method provided in the embodiment of the present application and present the natural language processing results to the user (the presentation method may be, but is not limited to, display, saving, uploading to the cloud, etc.).

[0080] In one possible implementation, a user can open a natural language processing application installed on a terminal device and enter text. The natural language processing application can send the text to a server on the cloud side. The server on the cloud side processes the text using the method provided in an embodiment of the present application and transmits the natural language processing results back to the terminal device. The terminal device can present the natural language processing results to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).

[0081] In one possible implementation, a user can open a natural language processing application installed on a terminal device and enter text. The natural language processing application can send the text to a server on the cloud side. The server on the cloud side performs natural language processing on the text using the method provided in an embodiment of the present application, and transmits the natural language processing results back to the terminal device. The terminal device can present the natural language processing results to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).

[0082] Next, the natural language processing application in the embodiment of this application is introduced from the functional architecture and the product architecture that implements the functions.

[0083] Referring to FIG. 1B , FIG. 1B is a schematic diagram of the functional architecture of a natural language processing application in an embodiment of the present application:

[0084] In one possible implementation, as shown in FIG1B , a natural language processing application 102 may receive input parameters 101 (e.g., including text) and generate predicted text or text natural language processing results 103. The natural language processing application 102 may be executed on (for example) at least one computer system and include computer text that, when executed by one or more computers, causes the computers to execute a natural language model trained using the method provided in the embodiments of the present application.

[0085] Referring to FIG. 1C , FIG. 1C is a schematic diagram of the physical architecture for running a natural language processing application in an embodiment of the present application:

[0086] Referring to FIG1C , FIG1C shows a schematic diagram of a system architecture. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers (FIG1C illustrates one server as an example), and the server 200 may provide natural language processing services for one or more terminals.

[0087] Among them, a natural language processing application can be installed on the terminal 100, or a web page related to the natural language processing function can be opened. The above application and web page can provide an interface. The terminal 100 can receive the relevant parameters entered by the user on the natural language processing function interface and send the above parameters to the server 200. The server 200 can obtain the processing results based on the received parameters and return the processing results to the terminal 100.

[0088] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters by itself without the need for the cooperation of the server, and the embodiments of the present application are not limited to this.

[0089] Next, the product form of the terminal 100 in FIG1C is described;

[0090] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.

[0091] FIG1D shows a schematic diagram of an optional hardware structure of the terminal 100 .

[0092] 1D , the terminal 100 may include components such as a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, and a power supply 190. Those skilled in the art will appreciate that FIG1D is merely an example of a terminal or multi-function device and does not limit the terminal or multi-function device. The terminal or multi-function device may include more or fewer components than shown, or may combine certain components or have different components.

[0093] The input unit 130 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the portable multifunction device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can detect user touch operations on or near it (for example, operations performed on or near the touch screen using a finger, joint, stylus, or any other suitable object) and drive corresponding connected devices according to pre-set programs. The touch screen can detect user touch actions on the touch screen, convert the touch actions into touch signals and transmit them to the processor 170. It can also receive and execute commands sent by the processor 170; the touch signals include at least touch point coordinate information. The touch screen 131 provides an input interface and an output interface between the terminal 100 and the user. Touch screens can be implemented using various types, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 131, the input unit 130 may also include other input devices. Specifically, the other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control button 132 , a switch button 133 , etc.), a trackball, a mouse, a joystick, and the like.

[0094] The input device 132 may receive input text, processing instructions, and the like.

[0095] The display unit 140 may be used to display information input by the user or provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In an embodiment of the present application, the display unit 140 may be used to display an interface of a natural language processing application, etc.

[0096] Memory 120 can be used to store instructions and data. It primarily includes an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files and text. The instruction storage area can store software units such as the operating system, applications, and instructions required for at least one function, or subsets or extensions thereof. It may also include non-volatile random access memory (RAM). It provides processor 170 with management functions for the hardware, software, and data resources within the computing and processing device, supporting control software and applications. It is also used to store multimedia files and running programs and applications.

[0097] The processor 170 is the control center of the terminal 100. It connects all components of the terminal 100 using various interfaces and circuits. By executing instructions stored in the memory 120 and accessing data stored therein, it executes various functions of the terminal 100 and processes data, thereby providing overall control of the terminal device. Optionally, the processor 170 may include one or more processing units. Preferably, the processor 170 may integrate an application processor and a modem processor, with the application processor primarily processing the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory may be implemented on a single chip; in other embodiments, they may be implemented on separate chips. The processor 170 may also generate corresponding operational control signals and send them to the corresponding components of the computing and processing device. It may also read and process data in the software, particularly the data and programs in the memory 120, to enable the various functional modules therein to perform their corresponding functions, thereby controlling the corresponding components to operate as instructed.

[0098] Among them, the memory 120 can be used to store software text related to the data processing method, the processor 170 can execute the steps of the chip's data processing method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to implement corresponding functions.

[0099] The RF unit 110 (optional) can be used to send and receive information or receive and send signals during a call. For example, after receiving downlink information from the base station, it is passed to the processor 170 for processing; in addition, it sends the designed uplink data to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF unit 110 can also communicate with network devices and other devices via wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0100] In this embodiment of the present application, the radio frequency unit 110 can send text to the server 200 and receive predicted text or natural language processing results sent by the server 200.

[0101] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.

[0102] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.

[0103] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to other devices for communication, or to connect a charger to charge the terminal 100 .

[0104] Although not shown, the terminal 100 may also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which are not described in detail here. Some or all of the methods described below may be applied to the terminal 100 shown in FIG1D .

[0105] Next, the product form of the server 200 in FIG1C is described;

[0106] FIG2 provides a schematic diagram of the structure of a server 200. As shown in FIG2, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other via the bus 201.

[0107] Bus 201 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, FIG2 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.

[0108] The processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0109] The memory 204 may include volatile memory, such as random access memory (RAM). The memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard drive (HDD), or solid state drive (SSD).

[0110] The memory 204 may be used to store software text related to the data processing method, and the processor 202 may execute the steps of the data processing method of the chip, and may also schedule other units to implement corresponding functions.

[0111] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0112] It should be understood that the steps related to the model reasoning process in the embodiments of this application involve AI-related operations. When performing AI operations, the instruction execution architecture of the terminal device and server is not limited to the processor-memory architecture described above. The system architecture provided in the embodiments of this application is described in detail below with reference to Figure 5.

[0113] FIG5 is a schematic diagram of the system architecture provided by an embodiment of the present application. As shown in FIG5 , the system architecture 500 includes an execution device 510 , a training device 520 , a database 530 , a client device 540 , a data storage system 550 , and a data acquisition system 560 .

[0114] The execution device 510 includes a calculation module 511, an I / O interface 512, a pre-processing module 513, and a post-processing module 514. The calculation module 511 may include the target model / rule 501, and the pre-processing module 513 and the post-processing module 514 are optional.

[0115] The execution device 510 may be a terminal device or a server that runs the above-mentioned natural language processing application.

[0116] The data collection device 560 is used to collect training samples. The training samples can be text, etc. After collecting the training samples, the data collection device 560 stores these training samples in the database 530.

[0117] The training device 520 can train the neural network to be trained (such as the language model in the embodiment of the present application, etc.) based on the training samples maintained in the database 530 to obtain the target model / rule 501.

[0118] It should be noted that, in actual applications, the training samples maintained in the database 530 may not all be collected by the data acquisition device 560, but may also be received from other devices. It should also be noted that the training device 520 may not train the target model / rule 501 entirely based on the training samples maintained in the database 530, but may also obtain training samples from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.

[0119] The target model / rule 501 obtained through training with the training device 520 can be applied to different systems or devices, such as the execution device 510 shown in FIG5 . The execution device 510 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, an in-vehicle terminal, etc., or a server, etc.

[0120] Specifically, the training device 520 may transfer the trained model to the execution device 510 .

[0121] In Figure 5, the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with external devices. The user can input data (such as text in the embodiments of the present application) to the I / O interface 512 through the client device 540.

[0122] Preprocessing module 513 and preprocessing module 514 are used to preprocess the input data received by I / O interface 512. It should be understood that preprocessing module 513 and preprocessing module 514 may be absent or only one preprocessing module may be present. If preprocessing module 513 and preprocessing module 514 are absent, computing module 511 may be used directly to process the input data.

[0123] When the execution device 510 preprocesses the input data, or when the computing module 511 of the execution device 510 performs calculations and other related processing, the execution device 510 can call the data, text, etc. in the data storage system 550 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 550.

[0124] Finally, the I / O interface 512 provides the processed results to the client device 540 and thus to the user.

[0125] In the scenario shown in FIG5 , the user can manually input data, and this "manual input data" can be operated through the interface provided by I / O interface 512. In another scenario, client device 540 can automatically send input data to I / O interface 512. If user authorization is required for client device 540 to automatically send input data, the user can set the corresponding permissions in client device 540. The user can view the results output by execution device 510 on client device 540, and the specific presentation form can be a display, sound, action, or other specific method. Client device 540 can also serve as a data acquisition terminal, collecting input data input into I / O interface 512 and output results from I / O interface 512 as new sample data and storing them in database 530. Of course, collection can also be performed without client device 540, and instead the I / O interface 512 directly stores the input data input into I / O interface 512 and output results from I / O interface 512 as new sample data in database 530.

[0126] It is worth noting that FIG5 is merely a schematic diagram of a system architecture provided by an embodiment of the present application, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in FIG5 , the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the execution device 510 can be deployed in the client device 540.

[0127] From the inference side of the model:

[0128] In the embodiment of the present application, the computing module 511 of the above-mentioned execution device 510 can obtain the text stored in the data storage system 550 to implement the steps related to the model reasoning process in the embodiment of the present application.

[0129] In an embodiment of the present application, the computing module 511 of the execution device 510 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0130] Specifically, the computing module 511 of the execution device 510 can be a hardware system with the function of executing instructions, and the steps related to the model reasoning process provided in the embodiment of the present application can be software text stored in the memory. The computing module 511 of the execution device 510 can obtain the software text from the memory and execute the obtained software text to implement the steps related to the model reasoning process provided in the embodiment of the present application.

[0131] It should be understood that the computing module 511 of the execution device 510 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to the model reasoning process provided in the embodiment of the present application can also be implemented by the hardware system that does not have the function of executing instructions in the computing module 511 of the execution device 510, which is not limited here.

[0132] From the training side of the model:

[0133] In an embodiment of the present application, the above-mentioned training device 520 can obtain the text stored in the memory (not shown in Figure 5, which can be integrated into the training device 520 or deployed separately from the training device 520) to implement the steps related to model training in the embodiment of the present application.

[0134] In an embodiment of the present application, the training device 520 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.

[0135] It should be understood that the training device 520 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to model training provided in the embodiments of the present application can also be implemented by the hardware system in the training device 520 that does not have the function of executing instructions, which is not limited here.

[0136] 2. Natural language processing cloud services provided by the server:

[0137] In one possible implementation, the server may provide a natural language processing service to the terminal through an application programming interface (API).

[0138] Among them, the terminal device can send relevant parameters (such as text) to the server through the API provided by the cloud. The server can obtain processing results based on the received parameters and return the processing results to the terminal.

[0139] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.

[0140] FIG6 shows the process of using a natural language processing cloud service provided by a cloud platform.

[0141] 1. Activate and purchase the natural language processing service.

[0142] 2. Users can download the software development kit (SDK) corresponding to the content review service. Usually, the cloud platform provides multiple development versions of the SDK for users to choose according to the requirements of the development environment, such as JAVA version SDK, Python version SDK, PHP version SDK, Android version SDK, etc.

[0143] 3. After the user downloads the corresponding version of the SDK to the local computer as needed, they import the SDK project into the local development environment, configure and debug it in the local development environment, and develop other functions in the local development environment to form an application that integrates natural language processing capabilities.

[0144] 4. When a natural language processing application is used and needs to perform natural language processing, it can trigger an API call for the natural language processing function. When the application triggers the natural language processing function, it initiates an API request to the running instance of the natural language processing function service in the cloud environment. The API request includes text, and the running instance in the cloud environment processes the text to obtain the processing results.

[0145] 5. The cloud environment returns the processing results to the application, thereby completing a natural language processing function service call.

[0146] 3. Word segmentation cloud service provided by the server:

[0147] In a possible implementation, the server may provide word segmentation services to the client through an application programming interface (API).

[0148] Among them, the terminal device can send relevant parameters (such as text) to the server through the API provided by the cloud. The server can obtain processing results based on the received parameters and return the processing results to the terminal.

[0149] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.

[0150] FIG7 shows the process of using a word segmentation cloud service provided by a cloud platform.

[0151] 1. Activate and purchase the word segmentation service.

[0152] 2. Users can download the software development kit (SDK) corresponding to the content review service. Usually, the cloud platform provides multiple development versions of the SDK for users to choose according to the requirements of the development environment, such as JAVA version SDK, Python version SDK, PHP version SDK, Android version SDK, etc.

[0153] 3. After the user downloads the corresponding version of the SDK to the local computer according to their needs, they import the SDK project into the local development environment, configure and debug it in the local development environment. The local development environment can also be used to develop other functions, forming an application that integrates word segmentation capabilities.

[0154] 4. When the word segmentation function is used, the application can trigger an API call when word segmentation is needed. When the application triggers the word segmentation function, it initiates an API request to the running instance of the word segmentation service in the cloud environment. The API request contains text, and the running instance in the cloud environment processes the text to obtain the processing result.

[0155] 5. The cloud environment returns the processing results to the application, thereby completing a service call of the word segmentation function.

[0156] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.

[0157] In order to better understand the solution of the embodiment of the present application, the following briefly introduces the possible application scenarios of the embodiment of the present application with reference to Figures 3 and 4.

[0158] Figure 3 shows a natural language processing system, which includes user devices and data processing equipment. User devices include intelligent terminals such as mobile phones, personal computers, or information processing centers. User devices are the initiators of natural language data processing, initiating requests such as language questions and answers or inquiries. Typically, users initiate requests through their user devices.

[0159] The aforementioned data processing devices can be devices or servers with data processing capabilities, such as cloud servers, network servers, application servers, and management servers. The data processing devices receive query statements, voice, text, and other information from smart terminals via interactive interfaces. They then use their memory and processors to perform language data processing, including machine learning, deep learning, search, reasoning, and decision-making, and then feed the results back to the user device. The memory in a data processing device is a general term encompassing both local storage and databases storing historical data. The databases can be located on the data processing device or on other network servers.

[0160] In the natural language processing system shown in Figure 3, the user device can receive user instructions. For example, the user device can receive a piece of text input by the user, and then initiate a request to the data processing device, so that the data processing device executes a natural language processing application (such as natural language generation, text classification, text reasoning, named entity recognition, translation, etc.) for the piece of text obtained by the user device, thereby obtaining the processing results of the corresponding natural language processing application for the piece of text (such as predicted word results, classification results, reasoning results, named entity recognition results, translation results, etc.).

[0161] In an embodiment of the present application, the user device can receive instructions from the user. For example, the user device can receive a piece of text (such as text) input by the user, and then initiate a request to the data processing device, so that the data processing device executes a natural language processing application for the piece of text obtained by the user device, thereby obtaining a processing result of the corresponding natural language processing application for the piece of text.

[0162] Figure 4 shows another natural language processing system. In Figure 4, the user device directly serves as a data processing device. The user device can directly receive input from the user and process it directly by the hardware of the user device itself. The specific process is similar to that of Figure 3. Please refer to the above description and will not be repeated here.

[0163] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.

[0164] (1) Neural Network

[0165] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:

[0166] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0167] (2) Transformer layer

[0168] The neural network includes an embedding layer and at least one transformer layer, and the at least one transformer layer can be N transformer layers (N is an integer greater than 0), wherein each transformer layer includes an attention layer, an add&norm layer, a feed forward layer, and an add&norm layer that are adjacent in sequence. In the embedding layer, the current input is embedded to obtain multiple embedding vectors; in the attention layer, P input vectors are obtained from the previous layer of the first transformer layer, and with any first input vector among the P input vectors as the center, based on the correlation between each input vector within a preset attention window and the first input vector, the intermediate vector corresponding to the first input vector is obtained, thereby determining the P intermediate vectors corresponding to the P input vectors; in the pooling layer, the P intermediate vectors are merged into Q output vectors, wherein the multiple output vectors obtained by the last transformer layer in the transformer layer are used as feature representations of the current input.

[0169] (3) Attention mechanism

[0170] The attention mechanism mimics the internal process of biological observation behavior, namely, a mechanism that aligns internal experience and external sensations to increase the observation precision of certain areas. It can quickly filter out high-value information from a large amount of information using limited attention resources. The attention mechanism can quickly extract important features from sparse data and is therefore widely used in natural language processing tasks, especially machine translation. The self-attention mechanism is an improvement on the attention mechanism, which reduces dependence on external information and is better at capturing the internal correlation of data or features. The essential idea of ​​the attention mechanism can be rewritten as the following formula:

[0171] Here, Lx = ||Source|| represents the length of the Source. This formula implies that the elements in the Source are imagined to consist of a series of data pairs. Given a Query element in the target, the similarity or correlation between the Query and each Key is calculated to obtain the weight coefficient for each Key's corresponding Value. The weighted sum of the Values ​​is then taken to obtain the final Attention value. Essentially, the Attention mechanism performs a weighted sum of the Values ​​of the Source elements, with the Query and Key used to calculate the weight coefficient for the corresponding Value. Conceptually, Attention can be understood as selectively filtering out a small amount of important information from a large amount of information and focusing on this important information, while ignoring the majority of less important information. This focusing process is reflected in the calculation of the weight coefficients: the larger the weight, the more focus is placed on the corresponding Value. In other words, the weight represents the importance of the information, while the Value represents the corresponding information. The self-attention mechanism can be understood as internal attention. The attention mechanism occurs between the Query element of the Target and all elements of the Source. The self-attention mechanism refers to the attention mechanism that occurs between the internal elements of the Source or the internal elements of the Target. It can also be understood as the attention calculation mechanism in the special case of Target = Source. The specific calculation process is the same, only the calculation object has changed.

[0172] (4) Natural language processing (NLP)

[0173] Natural language refers to human language, and natural language processing (NLP) is the processing of human language. Natural language processing is the process of systematically analyzing, understanding, and extracting information from text data in an intelligent and efficient manner. By using NLP and its components, we can manage very large amounts of text data, perform a large number of automated tasks, and solve a wide variety of problems, such as automatic summarization, machine translation (MT), named entity recognition (NER), relation extraction (RE), information extraction (IE), sentiment analysis, speech recognition, question answering, and topic segmentation.

[0174] (5) Pre-trained language model

[0175] A pretrained language model is a natural language sequence encoder that encodes each word in a natural language sequence into a vector representation for prediction. Its training consists of two phases. In the pre-training phase, the model is trained on a large amount of unsupervised text for language modeling tasks, thereby learning a word representation. In the fine-tuning phase, the model is initialized using the parameters learned in the pre-training phase and trained in a relatively short number of steps on downstream tasks such as text classification and sequence labeling. This allows the semantic information gained from pre-training to be successfully transferred to downstream tasks.

[0176] (6) Backpropagation algorithm

[0177] Convolutional neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during training, reducing the reconstruction error loss of the super-resolution model. Specifically, the forward propagation of the input signal to the output generates an error loss. This error loss information is then backpropagated to update the parameters of the initial super-resolution model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal super-resolution model parameters, such as the weight matrix.

[0178] (7) Loss function

[0179] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss as much as possible.

[0180] (8) Text sequence

[0181] A sequence of characters, such as a natural language text or a program code.

[0182] (9) Vocabulary

[0183] A list of words used by the language model. Each word is a string of arbitrary length without spaces, such as a complete natural language word, or a part of a complete word (also called a subword).

[0184] (10) Out of vocabulary

[0185] Words that do not appear in the vocabulary when the language model is performing inference.

[0186] (11) Text tokenization

[0187] The technology of segmenting text sequences according to certain rules aims to reduce the possibility of unregistered words appearing.

[0188] Existing natural language processing systems use Byte Pair Encoding (BPE) or SentencePiece algorithms to process input text, segmenting text strings into segments delimited by spaces according to a specific pattern. This is done to significantly reduce the occurrence of unrecognized words during inference and improve the generalization and error tolerance of the language model. Without segmentation, all words in the language must appear in the training text. Otherwise, missing words (such as spelling errors) will be treated as unrecognized words during inference, significantly impacting the language model's error tolerance and inference capabilities. Furthermore, without word segmentation, some words with the same origin but different word forms will be treated as independent by the language model. This prevents the shared semantics of these words from being fully learned, reducing the language model's generalization ability.

[0189] However, these algorithms focus more on building mathematical models to achieve a more optimal subword distribution, while (partially) ignoring the inconsistencies in word segmentation caused by the inherent characteristics of language. This means that the same word may have different segmentation results in different scenarios. A common solution is to introduce a text pre-processing (or post-processing) module before or after word segmentation. While this approach can effectively address the inconsistency issue, it often introduces noise, leading to lossy word segmentation. This means that the input text cannot be restored after pre-processing, word segmentation, reverse segmentation, and post-processing.

[0190] In order to solve the above problems, the present invention provides a data processing method. The data processing method of the present invention is described in detail below with reference to the accompanying drawings.

[0191] Referring to Figure 8, Figure 8 is a flow chart of a data processing method provided in an embodiment of the present application. As shown in Figure 8, a data processing method provided in an embodiment of the present application may include steps 801 to 803, and these steps are described in detail below.

[0192] 801. Get text.

[0193] In a possible implementation, the first text may be a text that needs to be segmented, or a text that needs to be subsequently processed by natural language processing. The text may also be referred to as a text sequence, including multiple characters and character strings.

[0194] 802. Insert a delimiter into the text, where the delimiter is a non-space character string and is used to identify a position where characters are separated in the text.

[0195] Existing techniques improve word segmentation consistency across different scenarios by inserting spaces as separators at corresponding locations in the text. While simple and convenient, these spaces are identical in form to naturally occurring spaces. This makes it impossible for post-processing algorithms to accurately remove the inserted spaces in some scenarios, resulting in an inability to fully restore the original text sequence and impaired word segmentation.

[0196] In an embodiment of the present application, a non-space character string (or a character string that is not entirely space) is defined as a delimiter. For example, a character string that starts with a single or multiple non-space characters and ends with a single or multiple space characters is used as a delimiter. Since non-space characters are not as common as spaces in natural language, the inserted delimiter can be accurately deleted during post-processing, thereby achieving lossless word segmentation.

[0197] Exemplarily, the delimiter may be a character corresponding to "\u2582" in the source code. The character may be a lower quarter block element.

[0198] In one possible implementation, the non-space character string mentioned above can be added as a delimiter in the word segmentation model. For example, characters in the text can be identified one by one to determine whether they meet the conditions for inserting a delimiter.

[0199] When segmenting a text based on a word segmentation model, a delimiter may be inserted into the text. For example, in one possible implementation, the delimiter may be inserted immediately after a punctuation mark included in the text. That is, the delimiter may be inserted immediately after a punctuation mark (including the beginning of a sentence) and immediately following a word.

[0200] For example, in one possible implementation, a separator can be inserted at a position in the text where different languages ​​are converted and not separated by spaces. That is, a separator can be inserted between two languages ​​at a position where different languages ​​are converted and not separated by spaces.

[0201] For example, in a possible implementation, a separator may be inserted at an adjacent position after the last space in the continuous spaces included in the text, that is, at the position of the continuous spaces, after the last space.

[0202] In a possible implementation, a string of characters that are the same as the delimiter may also exist in the text. However, this part of the string belongs to the text and is not used as a delimiter. During post-processing, this part of the characters should not be deleted. Therefore, the string of characters that are the same as the delimiter can be protected in the text, that is, a string embedding structure is constructed to protect the non-space character portion of the above delimiter that appears in the input text, forming a protected string. Furthermore, the text after the delimiter is inserted also includes: a protected string, which is used to identify the text in the text that is the same as the delimiter, and indicates that the text that is the same as the delimiter is not a delimiter.

[0203] 803. Perform word segmentation on the text after the separator is inserted to obtain a word segmentation result.

[0204] The word segmentation result may include multiple word segmentation units. By designing the above separators, the same word can obtain consistent word segmentation results in different scenarios.

[0205] In one possible implementation, a first processing result of the text may be obtained based on the word segmentation result through a language model. The first processing result includes multiple word segmentation units and the separators. To obtain a processing result corresponding to the text, the first processing result may be post-processed. The post-processing may include, but is not limited to, reverse word segmentation and removal of separators. That is, the first processing result may be reversely segmented and the separators may be removed to obtain a second processing result of the text.

[0206] In a possible implementation, the first processing result further includes the protection character string, and post-processing may further include: removing the protection character string from the first processing result.

[0207] More specifically, please refer to Figure 9, which is a process diagram of an embodiment of the present application. Figure 9 shows the insertion process of separators and protection strings, word segmentation and post-processing process (including inverse word segmentation, separator removal and removal of string protection structure).

[0208] 10 , which is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. As shown in FIG10 , a data processing device 1000 provided in an embodiment of the present application includes:

[0209] Acquisition module 1001, used to acquire text;

[0210] The specific description of the acquisition module 1001 can refer to the introduction of step 801 in the above embodiment, which will not be repeated here.

[0211] The processing module 1002 is used to insert a delimiter into the text, where the delimiter is a non-space character string and is used to identify the position where the characters in the text are separated; and perform word segmentation on the text after the delimiter is inserted to obtain a word segmentation result.

[0212] The specific description of the processing module 1002 can refer to the introduction of step 802 and step 803 in the above embodiment, which will not be repeated here.

[0213] In a possible implementation, the separator is a character corresponding to the source code "\u2582".

[0214] In a possible implementation, the processing module 1002 is specifically configured to:

[0215] Inserting a separator into the text based on at least one of the following means:

[0216] inserting a separator at an adjacent position after a punctuation mark included in the text;

[0217] Inserting separators at locations in the text where different languages ​​are converted and not separated by spaces;

[0218] A delimiter is inserted at a position immediately after the last of the consecutive spaces included in the text.

[0219] In a possible implementation, the processing module 1002 is further configured to:

[0220] According to the word segmentation result, a first processing result of the text is obtained through a language model.

[0221] In a possible implementation, the first processing result includes a plurality of word segmentation units and the separator; the processing module 1002 is further configured to:

[0222] Reverse word segmentation is performed on the first processing result, and the separator is removed to obtain a second processing result of the text.

[0223] In a possible implementation, the text after the separator is inserted further includes: a protection string, wherein the protection string is used to identify text in the text that has the same characters as the separator and indicates that the text that has the same characters as the separator is not a separator.

[0224] In a possible implementation, the first processing result further includes the protected character string; and the processing module 1002 is further configured to:

[0225] The protected character string is removed from the first processing result.

[0226] Next, an execution device provided in an embodiment of the present application is introduced. Please refer to Figure 11. Figure 11 is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1100 can be specifically manifested as a virtual reality VR device, a mobile phone, a tablet, a laptop computer, a smart wearable device, a monitoring data processing device or a server, etc., which is not limited here. Specifically, the execution device 1100 includes: a receiver 1101, a transmitter 1102, a processor 1103 and a memory 1104 (wherein the number of processors 1103 in the execution device 1100 can be one or more, and Figure 11 takes one processor as an example), wherein the processor 1103 may include an application processor 11031 and a communication processor 11032. In some embodiments of the present application, the receiver 1101, the transmitter 1102, the processor 1103 and the memory 1104 may be connected via a bus or other means.

[0227] The memory 1104 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1103. A portion of the memory 1104 may also include non-volatile random access memory (NVRAM). The memory 1104 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0228] Processor 1103 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all of these buses are referred to as a bus system in the figure.

[0229] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1103. Processor 1103 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1103. The above processor 1103 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1103 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1104. Processor 1103 reads information from memory 1104 and, in conjunction with its hardware, completes the steps involved in the model inference process in the above method.

[0230] Receiver 1101 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1102 can be used to output digital or character information through the first interface. Transmitter 1102 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1102 can also include a display device such as a display screen.

[0231] The present application also provides a training device. Please refer to FIG. 12 , which is a schematic diagram of the structure of a training device provided by an embodiment of the present application. Specifically, the training device 1200 is implemented by one or more servers. The training device 1200 may vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1212 (e.g., one or more processors), a memory 1232, and one or more storage media 1230 (e.g., one or more mass storage devices) storing application programs 1242 or data 1244. The memory 1232 and storage medium 1230 may be either ephemeral or persistent storage. The program stored in the storage medium 1230 may include one or more modules (not shown), each of which may include a series of instruction operations on the training device. Furthermore, the CPU 1212 may be configured to communicate with the storage medium 1230 to execute the series of instruction operations in the storage medium 1230 on the training device 1200.

[0232] The training device 1200 may also include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input and output interfaces 1258; or, one or more operating systems 1241, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0233] In the embodiment of the present application, the central processing unit 1212 is used to execute actions related to model training in the above embodiment.

[0234] An embodiment of the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.

[0235] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.

[0236] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the data processing method described in the above embodiment, or so that the chip in the training device executes the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0237] Specifically, see Figure 13 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip may be a neural network processor (NPU) 1300. NPU 1300 is mounted on a host CPU (host CPU) as a coprocessor, with tasks assigned by the host CPU. The core of the NPU is arithmetic circuit 1303, which is controlled by controller 1304 to extract matrix data from memory and perform multiplication operations.

[0238] In some implementations, the arithmetic circuit 1303 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1303 is a two-dimensional systolic array. The arithmetic circuit 1303 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1303 is a general-purpose matrix processor.

[0239] For example, assume there are input matrix A, weight matrix B, and output matrix C. The computation circuit retrieves the corresponding data of matrix B from weight memory 1302 and caches it on each PE in the computation circuit. The computation circuit then retrieves the data of matrix A from input memory 1301 and performs a matrix operation on it with matrix B. The partial or final matrix result is stored in accumulator 1308.

[0240] Unified memory 1306 is used to store input and output data. Weight data is directly transferred to weight memory 1302 through the Direct Memory Access Controller (DMAC) 1305. Input data is also transferred to unified memory 1306 through the DMAC.

[0241] BIU stands for Bus Interface Unit 1310 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1309 .

[0242] The bus interface unit 1310 (BIU) is used for the instruction fetch memory 1309 to obtain instructions from the external memory, and is also used for the storage unit access controller 1305 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0243] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1306 or to move weight data to the weight memory 1302 or to move input data to the input memory 1301.

[0244] The vector calculation unit 1307 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1303, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0245] In some implementations, the vector calculation unit 1307 can store the processed output vector in the unified memory 1306. For example, the vector calculation unit 1307 can apply a linear function or a nonlinear function to the output of the operation circuit 1303, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1307 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1303, for example, for use in subsequent layers in a neural network.

[0246] An instruction fetch buffer 1309 connected to the controller 1304 is used to store instructions used by the controller 1304;

[0247] Unified memory 1306, input memory 1301, weight memory 1302, and instruction fetch memory 1309 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0248] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0249] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0250] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0251] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0252] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

Claims

1. A data processing method, characterized in that: The method comprises: Get text; Inserting a separator into the text, wherein the separator is a non-space character string and is used to identify a position where a character is separated in the text; The text after the separator is inserted is segmented to obtain a segmentation result.

2. The method according to claim 1, characterized in that: The delimiter is a character string that starts with a non-space character string and ends with at least one space character string.

3. The method according to claim 1 or 2, characterized in that: The separator is the character corresponding to the source code "\u2582".

4. The method according to any one of claims 1 to 3, characterized in that: The inserting a separator into the text comprises: Inserting a separator into the text based on at least one of the following methods: inserting a separator at an adjacent position after a punctuation mark included in the text; Insert separators at locations in the text where different languages ​​are converted and not separated by spaces; A separator is inserted at an adjacent position after the last space among consecutive spaces included in the text.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: According to the word segmentation result, a first processing result of the text is obtained through a language model.

6. The method according to any one of claims 1 to 5, characterized in that: The first processing result includes a plurality of word segmentation units and the separator; the method further includes: performing reverse word segmentation on the first processing result and removing the separator to obtain a second processing result of the text; or, The method further includes: performing reverse word segmentation on the word segmentation result and removing the separator to obtain a second processing result of the text.

7. The method according to any one of claims 1 to 6, characterized in that: The text after the separator is inserted further includes: a protection string, wherein the protection string is used to identify the text in the text that has the same characters as the separator, and to indicate that the text that has the same characters as the separator is not a separator.

8. The method according to claim 6 or 7, characterized in that: The first processing result also includes the protected character string; and the method further includes: The protection character string in the first processing result is removed.

9. A data processing device, characterized in that: The device comprises: Acquisition module, used to acquire text; The processing module is used to insert a separator into the text, wherein the separator is a non-space character string and is used to identify the position of the character separation in the text; and perform word segmentation on the text after the separator is inserted to obtain a word segmentation result.

10. The device according to claim 9, characterized in that The delimiter is a character string that starts with a non-space character string and ends with at least one space character string.

11. The device according to claim 9 or 10, characterized in that The separator is the character corresponding to the source code "\u2582".

12. The device according to any one of claims 9 to 11, characterized in that: The processing module is specifically used for: Inserting a separator into the text based on at least one of the following means: inserting a separator at an adjacent position after a punctuation mark included in the text; Insert separators at locations in the text where different languages ​​are converted and not separated by spaces; A separator is inserted at an adjacent position after the last space among consecutive spaces included in the text.

13. The device according to any one of claims 9 to 12, characterized in that: The processing module is further used for: According to the word segmentation result, a first processing result of the text is obtained through a language model.

14. The device according to any one of claims 9 to 13, characterized in that: The first processing result includes a plurality of word segmentation units and the separator; the processing module is further used to: perform reverse word segmentation on the first processing result and remove the separator to obtain a second processing result of the text; or, The processing module is further used to: perform reverse word segmentation on the word segmentation result and remove the separator to obtain a second processing result of the text.

15. The device according to any one of claims 9 to 14, characterized in that: The text after the separator is inserted further includes: a protection string, wherein the protection string is used to identify the text in the text that has the same characters as the separator, and to indicate that the text that has the same characters as the separator is not a separator.

16. The device according to claim 14 or 15, characterized in that The first processing result also includes the protected character string; and the processing module is further used to: The protection character string in the first processing result is removed.

17. A computer storage medium, characterized in that: The computer storage medium stores one or more instructions which, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 8.

18. A computer program product, characterized in that The method comprises computer-readable instructions, and when the computer-readable instructions are executed on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 8.

19. A system comprising at least one processor and at least one memory; the processor and the memory are connected via a communication bus and communicate with each other; The at least one memory is used to store text; The at least one processor is used to execute the text to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Text data processing method, device and equipment

    CN110032730A

  • Word segmentation method and server

    CN110162794A

  • Text processing method and device, model training method and device, equipment and storage medium

    CN116306527A

  • Data processing method and device

    CN117892700A

  • Storing a data structure

    US20230110803A1