Data processing method, device, equipment and storage medium

By performing word segmentation on the input text fragments of the large model, screening out high-frequency words and storing their embedding results, establishing a mapping relationship, and directly calling the embedding results of high-frequency words, the problems of large model's high computing resource consumption and low efficiency are solved, and computing resource savings and efficiency improvements are achieved.

CN120336510BActive Publication Date: 2025-09-19INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510797913.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-19
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Large models require a lot of computing resources during training and inference, resulting in high computing resource consumption and low computing efficiency.

Method used

By performing word segmentation on the input text fragment, determining the word set and its frequency information, screening out the target word whose frequency meets the preset conditions, and storing its corresponding embedding results in the preset storage area, a mapping relationship is established between the word and the storage address of the embedding result, so that the embedding results of the high-frequency word can be directly called, omitting repeated embedding layer calculations.

Benefits of technology

It saves computing resources and improves computing efficiency, especially on devices with limited computing power but large memory, which can significantly improve computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336510B_ABST
    Figure CN120336510B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method, device, equipment and storage medium, which relates to the field of large model technology, including: based on the word segmentation processing of the input first text fragment, determining the frequency information of the first word unit set and each first word unit in the first word unit set; taking the first word unit in the first word unit set whose frequency information meets the preset conditions as the target word unit; obtaining the embedding result corresponding to the target word unit, and storing the embedding result corresponding to the target word unit in a preset storage area; establishing a mapping relationship between the target word unit and the storage address of the corresponding embedding result. In this way, the word units that appear frequently are screened out according to the statistics of the frequency information, and the embedding results corresponding to the target word units are stored in the preset storage area, so that when the text fragment is subsequently inferred, the embedding results of the target word units that appear frequently can be directly called, and the embedding layer calculation process of the target word units that appear frequently is omitted, thereby saving computing resources and improving computing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model technology, and in particular to a data processing method, apparatus, device and storage medium. Background Art

[0002] Large models, as machine learning models with large parameters and complex computational structures, are capable of handling complex tasks and data, and have broad applications across various fields. However, the application of large models relies on powerful computing capabilities, requiring significant computational resources during training and inference. Larger models require more computing resources and take longer to compute, leading to high resource consumption and low computational efficiency. Summary of the Invention

[0003] The present invention provides a data processing method, apparatus, device and storage medium to at least solve the problems of large computing resource consumption and low computing efficiency existing in large models in related technologies.

[0004] The present invention provides a data processing method, comprising:

[0005] Determining a first word-gram set and frequency information of each first word-gram in the first word-gram set based on word segmentation processing of the input first text segment;

[0006] The first word-gram whose frequency information meets the preset conditions in the first word-gram set is used as the target word-gram;

[0007] Obtain the embedding result corresponding to the target word unit, and store the embedding result corresponding to the target word unit in a preset storage area;

[0008] A mapping relationship is established between the target word and the storage address of the corresponding embedding result, and the mapping relationship is used to search and call the embedding result corresponding to the target word.

[0009] The present invention also provides a data processing device, comprising:

[0010] a frequency information determination module, configured to determine the frequency information of the first word-gram set and each first word-gram in the first word-gram set based on word segmentation processing of the input first text segment;

[0011] a target word-unit determination module, configured to take a first word-unit in the first word-unit set whose frequency information meets a preset condition as a target word-unit;

[0012] The embedding result storage module is used to obtain the embedding result corresponding to the target word and store the embedding result corresponding to the target word in a preset storage area;

[0013] The mapping relationship establishment module is used to establish a mapping relationship between the target word and the storage address of the corresponding embedding result. The mapping relationship is used to search and call the embedding result corresponding to the target word.

[0014] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any one of the above-mentioned data processing methods when executing the computer program.

[0015] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data processing methods are implemented.

[0016] The present invention also provides a computer program product, comprising a computer program, which implements the steps of any of the above data processing methods when executed by a processor.

[0017] Through the present invention, word segmentation is performed on the first input text fragment, and the frequency information of the first word obtained by the word segmentation is counted, and the first word whose frequency information meets the preset conditions is used as the target word, so that the frequently appearing words are screened out according to the statistics of the frequency information; at the same time, the embedding result corresponding to the target word is stored in a preset storage area, and a mapping relationship is established between the target word and the storage address of the corresponding embedding result, so that when the text fragment is subsequently inferred, the embedding result of the frequently appearing target word can be directly called, and the embedding layer calculation process of the frequently appearing target word is omitted, thereby saving computing resources and improving computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 is a flow chart of a data processing method according to an embodiment of the present invention;

[0020] Figure 2 is a schematic diagram of a preset storage area in a data processing method according to an embodiment of the present invention;

[0021] Figure 3 is a flow chart of another data processing method according to an embodiment of the present invention;

[0022] Figure 4 is a flowchart of another data processing method according to an embodiment of the present invention;

[0023] Figure 5 is a schematic diagram of a specific embodiment of the data processing method according to an embodiment of the present invention;

[0024] Figure 6 is a schematic diagram of another specific embodiment of the data processing method according to an embodiment of the present invention;

[0025] Figure 7 is a structural block diagram of a data processing device according to an embodiment of the present invention;

[0026] Figure 8 FIG. 4 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0030] Large models, as machine learning models with large parameters and complex computational structures, are capable of handling complex tasks and data, and have broad applications across various fields. However, the application of large models relies on powerful computing capabilities, requiring significant computational resources during training and inference. Larger models require more computing resources and take longer to compute, leading to high resource consumption and low computational efficiency.

[0031] Specifically, the inference process of a large model is actually the process of predicting the next token based on the previously input or generated token. Therefore, each prediction requires the completion of the forward calculation of the current last token. The currently generated token is first processed by the tokenizer into an integer token number, and then passes through the embedding layer, turning the integer into a tensor. The tensor is then calculated through each layer of the transformer structure, where each layer undergoes attention and multi-layer perception (MLP) processing, and finally outputs the calculation result. In fact, before the transformer calculation, the embedding result obtained by the embedding operation for the same token is the same. However, during the inference process of the large model, when a token appears again, it will still undergo repeated embedding layer calculations during the inference process, resulting in a waste of computing resources.

[0032] In related technologies, there is a solution that uses the Multilingual BERT model to prune the embedding layer of low-frequency words. This solution uses the BERT model's vocabulary to segment the target corpus during inference, calculates word frequency, and prunes the embedding layer of particularly low-frequency words, thereby reducing the number of parameters in the embedding layer. However, this method only reduces the number of model parameters and memory consumption, but does not reduce the computational operations in the inference process, nor can it improve the computational efficiency of the inference process of large models.

[0033] Based on this, the present invention provides a data processing method, which determines a first word set and the frequency information of each first word in the first word set based on the word segmentation processing of the input first text fragment; takes the first word in the first word set whose frequency information meets the preset conditions as the target word; obtains the embedding result corresponding to the target word, and stores the embedding result corresponding to the target word in a preset storage area; establishes a mapping relationship between the target word and the storage address of the corresponding embedding result, and the mapping relationship is used to search and call the embedding result corresponding to the target word. Thus, the input first text fragment is segmented, and the frequency information of the first word obtained by the word segmentation is counted, and the first word whose frequency information meets the preset conditions is used as the target word, so that the word that appears frequently is screened out according to the statistics of the frequency information; at the same time, the embedding result corresponding to the target word is stored in the preset storage area, and a mapping relationship is established between the target word and the storage address of the corresponding embedding result, so that when the text fragment is subsequently inferred, the embedding result of the target word that appears frequently can be directly called, and the embedding layer calculation process of the target word that appears frequently is omitted, thereby saving computing resources and improving computing efficiency.

[0034] The data processing method provided by the embodiments of the present invention can be applied not only to large models, but also to other models that require embedded computations. This method can significantly improve computational efficiency, particularly for devices with limited computing power but large memory. Furthermore, if the applied model structure does not require rotational position encoding calculations, the final computational results of the target word can be directly stored in a preset storage area, thereby directly omitting the model inference computation process and further improving computational efficiency.

[0035] This embodiment provides a data processing method that can be used for various models with embedded calculations. Figure 1 FIG. 1 is a flow chart of a data processing method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0036] Step S101 : determining a first word-gram set and frequency information of each first word-gram in the first word-gram set based on word segmentation processing of an input first text segment.

[0037] In an embodiment of the present invention, word segmentation processing is performed on the first text segment input into the model. Any word segmentation method can be used for word segmentation processing, and no specific limitation is imposed herein. Tokens obtained from the word segmentation processing of the first text segment are used as first tokens to form a first token set. The number of occurrences of each first token in the first text segment is counted to obtain frequency information for each first token.

[0038] Step S102 : taking the first word-gram in the first word-gram set whose frequency information meets a preset condition as a target word-gram.

[0039] In an embodiment of the present invention, the target word is a word that appears frequently during the reasoning process, and the embedding result corresponding to the target word will be stored in a preset storage area for easy search and call; wherein, the embedding result corresponding to the target word refers to the calculation result output after the target word is calculated by the embedding layer.

[0040] In an optional embodiment, the preset condition includes: the frequency information reaching a preset threshold; and / or the ratio of the frequency information to the number of first word-grams in the first word-gram set reaching a preset ratio. The preset threshold may be user-set or a pre-set fixed value.

[0041] Step S103: Obtain the embedding result corresponding to the target word-gram, and store the embedding result corresponding to the target word-gram in a preset storage area.

[0042] In this embodiment of the present invention, after the target word is calculated by the embedding layer, the calculation result output by the embedding layer is obtained to obtain the embedding result corresponding to the target word, and the embedding result corresponding to the target word is stored in a preset storage area. The preset storage area is a storage area pre-allocated in the memory for storing the embedding result.

[0043] In an optional embodiment, the preset storage area includes at least one storage area, and the at least one storage area can be dispersed in the memory, so that the storage area can be allocated at various locations in the memory and memory fragments can be fully utilized. Figure 2 Schematic diagram of a preset storage area in a data processing method according to an embodiment of the present invention. Figure 2 As shown, the preset storage area can be divided into multiple storage areas, which are dispersed in the unused memory according to the size of the memory fragments. Storage area No. 1 is located in one memory fragment, and memory areas No. 2 and No. 3 are located in another memory fragment, thereby making full use of the memory fragments.

[0044] Accordingly, the preset storage area can be divided and allocated in the following manner: First, based on the storage setting for the target word, determine the number of storage areas and the number of target word corresponding to the storage area, wherein the storage setting for the target word includes the setting of the preset number of target word and the setting of the target word group storage; the setting of the preset number of target word limits the number of high-frequency word stored, and the setting of the target word group storage limits the number of embedding results corresponding to the target word stored in a storage area. The number of storage areas can be obtained by dividing the two. For example, the preset number of target word can be set to 40, and 4 target words are set as a group, then 10 storage areas need to be divided. Then, based on the number of target word corresponding to the storage area and the storage space requirement corresponding to the target word, determine the size of the storage area; for example, 4 target words are set as a group, each F32 data format occupies 4 bytes, then the size of a storage area is 4*number of embedding layers*4Bytes.

[0045] Step S104: establishing a mapping relationship between the target word and the storage address of the corresponding embedding result.

[0046] In this embodiment of the present invention, the mapping relationship between the target word and the storage address of the corresponding embedding result is used to search and call the embedding result corresponding to the target word. During the subsequent reasoning process for a newly input text fragment, the mapping relationship can be first searched to determine whether the word obtained by word segmentation is the target word. If it is the target word, the mapping relationship can be used to determine the storage address of the corresponding embedding result. The data at the corresponding storage address can be read from the memory to obtain the embedding result, thereby omitting the calculation process of the embedding layer and improving computational efficiency.

[0047] In an optional embodiment, a storage list can be pre-constructed to record the storage address corresponding to the preset storage area and the target word corresponding to the embedding result stored in the storage address. When the embedding result corresponding to the target word is stored in the preset storage area, the target word is filled in the storage list at the location corresponding to the storage address where the embedding result corresponding to the target word is stored, thereby establishing a mapping relationship between the target word and the storage address of the corresponding embedding result. In the subsequent reasoning process, the embedding result corresponding to the target word can be quickly located by directly searching the storage list.

[0048] The data processing method provided by the embodiment of the present invention determines the frequency information of the first word element set and each first word element in the first word element set based on the word segmentation processing of the input first text segment; takes the first word element in the first word element set whose frequency information meets the preset conditions as the target word element; obtains the embedding result corresponding to the target word element, and stores the embedding result corresponding to the target word element in a preset storage area; establishes a mapping relationship between the target word element and the storage address of the corresponding embedding result, and the mapping relationship is used to search and call the embedding result corresponding to the target word element. Thus, the input first text segmentation processing is performed, and the frequency information of the first word element obtained by the word segmentation is counted, and the first word element whose frequency information meets the preset conditions is used as the target word element, thereby filtering out the word elements that appear frequently based on the statistics of the frequency information; at the same time, the embedding result corresponding to the target word element is stored in the preset storage area, and a mapping relationship is established between the target word element and the storage address of the corresponding embedding result, so that when the text segment is subsequently inferred, the embedding result of the target word element that appears frequently can be directly called, omitting the embedding layer calculation process of the target word element that appears frequently, thereby saving computing resources and improving computing efficiency.

[0049] This embodiment provides a data processing method that can be used for various models with embedded calculations. Figure 3 FIG. 1 is a flow chart of another data processing method according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0050] Step 301 : Based on word segmentation processing of an input first text segment, determine a first word-gram set and frequency information of each first word-gram in the first word-gram set.

[0051] In an embodiment of the present invention, when the preset storage area for storing embedding results is not full or the resource utilization rate does not reach the preset utilization rate, the text segment of the input model is used as the first text segment, and the first word element obtained by the word segmentation processing of the first text segment is accumulated and counted. Accordingly, step S301 includes:

[0052] Step S3011 : performing word segmentation processing on the input first text segment to determine a first word unit contained in the first text segment to update a first word unit set.

[0053] In an embodiment of the present invention, when the preset storage area for storing embedding results is not yet full or the resource utilization rate has not reached the preset utilization rate, each time a new first text fragment is input into the model, the input first text fragment is segmented to obtain the first word element contained in the first text fragment, and the first word element is compared with the first word element contained in the first word element set to determine the first word element in the first word element contained in the first text fragment that is not stored in the first word element set, that is, the first word element that appears for the first time, and update it to the first word element set.

[0054] Step S3012: Count the number of occurrences of the first word contained in the first text segment in the first text segment to obtain frequency information to be updated.

[0055] In the embodiment of the present invention, the number of occurrences of each first word contained in the input first text segment is counted respectively, and the number of occurrences of the first word in the first text segment is used as the frequency information to be updated corresponding to the first word.

[0056] Step S3013: Obtain frequency information of the first word contained in the first text segment, and update the frequency information based on the frequency information to be updated.

[0057] In an embodiment of the present invention, the frequency information of the first word contained in the first text segment is obtained. The frequency information is a statistical count of the total number of times the first word appears in all the first text segments input at historical moments. On the basis of the frequency information, the frequency information to be updated is superimposed to update the frequency information.

[0058] Step S302: The first word-gram in the first word-gram set whose frequency information meets the preset conditions is used as the target word-gram. Figure 1 Step S102 of the illustrated embodiment will not be described in detail here.

[0059] Step S303: Obtain the embedding result corresponding to the target word, and store the embedding result corresponding to the target word in a preset storage area. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0060] Step S304: Establish a mapping relationship between the target word and the storage address of the corresponding embedding result. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0061] The data processing method provided by an embodiment of the present invention performs word segmentation processing on an input first text fragment, determines the first word element contained in the first text fragment to update the first word element set, counts the number of occurrences of the first word element contained in the first text fragment in the first text fragment to obtain frequency information to be updated, obtains the frequency information of the first word element contained in the first text fragment, and updates the frequency information based on the frequency information to be updated, thereby continuously updating the first word element set and the frequency information of each first word element based on the newly input first text fragment, thereby ensuring the accuracy and reliability of the target word element determined based on the frequency information.

[0062] This embodiment provides a data processing method that can be used for various models with embedded calculations. Figure 4 is a flow chart of another data processing method according to an embodiment of the present invention, such as Figure 4 As shown, the process includes the following steps:

[0063] Step S401: Based on the word segmentation of the input first text segment, determine the first word set and the frequency information of each first word in the first word set. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.

[0064] Step S402: The first word-gram in the first word-gram set whose frequency information meets the preset conditions is used as the target word-gram. Figure 1 Step S102 of the illustrated embodiment will not be described in detail here.

[0065] Step S403: Obtain the embedding result corresponding to the target word, and store the embedding result corresponding to the target word in a preset storage area. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0066] Step S404: Establish a mapping relationship between the target word and the storage address of the corresponding embedding result. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0067] Step S405 , if the resource utilization of the preset storage area reaches the preset utilization, then within each data evaluation cycle, based on the word segmentation processing of the input second text segment, determine the second word set and the frequency information of each second word in the second word set.

[0068] In an embodiment of the present invention, if the resource utilization rate of the preset storage area reaches the preset utilization rate, it indicates that the preset storage area can no longer store the embedding results corresponding to more target word units. At this time, a storage competition optimization strategy is adopted to ensure that the embedding results corresponding to the most frequently appearing word units are always stored in the preset storage area, thereby maximizing the utilization efficiency of storage resources.

[0069] In an embodiment of the present invention, in the storage contention optimization strategy, within each data evaluation cycle, based on the word segmentation processing of the input second text segment, a second word-unit set and the frequency information of each second word-unit in the second word-unit set are determined. After the storage contention optimization strategy is enabled, that is, when the resource utilization rate of the preset storage area reaches the preset utilization rate, the text segments input within each data evaluation cycle are respectively regarded as second text segments, that is, the text segments are calculated and counted separately within each data evaluation cycle; accordingly, the second word-unit set is used to count the second words within a data evaluation cycle.

[0070] In the embodiment of the present invention, in each data evaluation cycle, the word segmentation process of the input second text segment is performed to determine the second word set and the frequency information of each second word in the second word set. For details, please refer to Figure 3 Step S301 in the illustrated embodiment will not be described in detail here.

[0071] In an optional embodiment, the data evaluation cycle can be determined based on the rounds of reasoning tasks, that is, the model will complete a preset round of reasoning tasks within one data evaluation cycle.

[0072] Step S406: Update the target word-gram based on the frequency information of the second word-gram in the second word-gram set.

[0073] In the embodiment of the present invention, based on the frequency information of the second word-grams in the second word-gram set, the second word-grams that appear frequently in the second word-gram set are determined and used as the target word-grams to update the target word-grams.

[0074] In an optional embodiment, the following method can be used to update the target word based on the frequency information of the second word in the second word set: based on the frequency information, the second word in the second word set is sorted to obtain the arrangement order of the second word in the second word set; wherein, the second word in the second word set is sorted according to the size of the frequency information, and the second word is arranged in an arrangement order from large to small or from small to large according to the frequency information. According to the arrangement order, a preset number of second word-grams are extracted from the second word set as target word-grams; wherein, if the second word-grams are arranged in an arrangement order from large to small according to the frequency information, a preset number of second word-grams are extracted from the front to back according to the arrangement order as target word-grams; if the second word-grams are arranged in an arrangement order from small to large according to the frequency information, a preset number of second word-grams are extracted from the back to front according to the arrangement order as target word-grams, thereby extracting the preset number of second word-grams with the highest frequency in the second word set, ensuring that the embedding results corresponding to the most frequently occurring word-grams are stored.

[0075] In an optional embodiment, a second word-gram whose frequency information meets a preset condition can also be selected from the second word-gram set as a target word-gram. If the number of target word-grams selected from the second word-gram set does not reach the upper limit of the number of target word-grams, that is, the preset number, the target word-grams corresponding to the embedding results stored in the preset storage area are filtered according to the frequency information, and target word-grams with lower frequency information are filtered out. The number of filtered target word-grams is consistent with the number of target word-grams selected from the second word-gram set, and the filtered target word-grams are cleared to release storage space for the target word-grams selected from the second word-gram set.

[0076] Step S407: Based on the updated embedding result corresponding to the target word, the embedding result stored in the preset storage area is updated.

[0077] In an embodiment of the present invention, considering that the target word corresponding to the embedding result already stored in the preset storage area may partially overlap with the updated target word, the embedding results stored in the preset storage area for the overlapping target word are retained, and the storage space in the preset storage area for the embedding results corresponding to the non-overlapping target word is released.

[0078] Specifically, first, the embedding results stored in the preset storage area are compared with the updated target word-gram to determine the word-gram to be stored and the storage area to be replaced. The word-gram to be stored refers to the word-gram whose corresponding embedding result in the target word-gram is not stored in the preset storage area, and the storage area to be replaced refers to the storage area where the word-gram corresponding to the embedding result stored in the preset storage area is no longer the target word-gram. The embedding results stored in the storage area to be stored need to be cleared to free up storage space. Then, the embedding results stored in the storage area to be replaced are cleared, and the embedding results corresponding to the word-gram to be stored are stored in the storage area to be replaced.

[0079] The data processing method provided by an embodiment of the present invention, when the resource utilization rate of the preset storage area reaches the preset utilization rate, determines the second word set and the frequency information of each second word in the second word set based on the word segmentation processing of the input second text fragment in each data evaluation cycle, and updates the target word based on the frequency information of the second word in the second word set, thereby determining the second word set and the frequency information of each second word in the second word set according to the data evaluation cycle, so as to be able to dynamically adjust the stored frequently occurring word elements in real time according to the reasoning task of the model; at the same time, based on the embedding result corresponding to the updated target word element, updates the embedding result stored in the preset storage area, thereby ensuring that the preset storage area always stores the embedding result corresponding to the most frequently occurring word element, so that the embedding result stored in the preset storage area can match the reasoning task of the actual model, ensuring that the embedding result stored in the preset storage area can play a maximum role, and improving the computational efficiency of the model.

[0080] As a specific implementation method, Figure 5 is a schematic diagram of a specific embodiment of the data processing method according to an embodiment of the present invention, Figure 5 The figure shows the process of processing word elements when the preset storage area is not full, that is, the resource utilization rate of the preset storage area does not reach the preset utilization rate. Figure 5 As shown in the figure, on the processor (CPU) side, it is determined whether the word unit obtained by word segmentation of the text fragment is the target word unit. On the graphics processing unit (GPU) side, the word unit is processed into an integer word unit number through the tokenizer component, and the word unit is converted from an integer to a tensor through the embedding layer to obtain the embedding result for the subsequent model inference process.

[0081] First, the CPU determines whether a word has been stored—that is, whether it has been identified as a target word. If so, the GPU retrieves the embedding result directly from a pre-set storage area for subsequent model inference. If not, the word's count is incremented by 1 to accumulate and update frequency information. After this update, the GPU determines whether the word meets the storage requirements and can be identified as a target word. If so, a storage address is allocated from the pre-set storage area. The GPU then processes the word into an integer number using the tokenizer component. The embedding layer converts the integer into a tensor, resulting in an embedding result. This result is then stored in the allocated storage address, and the model inference process continues. If not, the GPU processes the word into an integer number using the tokenizer component, converts the integer into a tensor, and the embedding layer converts the integer into a tensor, resulting in an embedding result, and the model inference process continues.

[0082] As a specific implementation method, Figure 6 is a schematic diagram of another specific embodiment of the data processing method according to an embodiment of the present invention, Figure 6 The figure shows the process of processing word elements when the preset storage area is full, that is, the resource utilization rate of the preset storage area reaches the preset utilization rate. Figure 6 As shown, within a data evaluation cycle, a specified round of reasoning tasks is completed, the frequency of occurrence of the word units appearing therein is counted and sorted, and the high-frequency word units, i.e., the target word units, are determined based on the sorting order. Then, a determination is made as to whether the word unit corresponding to the embedding result stored in the preset storage area is still the target word unit. If so, the corresponding embedding result stored in the preset storage area is retained. If not, the corresponding embedding result stored in the preset storage area is cleared to free up the preset storage area, and the embedding result corresponding to the unstored target word unit is stored in the preset storage area. This completes the determination of the target word unit and the storage of the embedding result within a data evaluation cycle.

[0083] The data processing method provided by the embodiment of the present invention determines the frequency information of the first word element set and each first word element in the first word element set based on the word segmentation processing of the input first text segment; takes the first word element in the first word element set whose frequency information meets the preset conditions as the target word element; obtains the embedding result corresponding to the target word element, and stores the embedding result corresponding to the target word element in a preset storage area; establishes a mapping relationship between the target word element and the storage address of the corresponding embedding result, and the mapping relationship is used to search and call the embedding result corresponding to the target word element. Thus, the input first text segmentation processing is performed, and the frequency information of the first word element obtained by the word segmentation is counted, and the first word element whose frequency information meets the preset conditions is used as the target word element, thereby filtering out the word elements that appear frequently based on the statistics of the frequency information; at the same time, the embedding result corresponding to the target word element is stored in the preset storage area, and a mapping relationship is established between the target word element and the storage address of the corresponding embedding result, so that when the text segment is subsequently inferred, the embedding result of the target word element that appears frequently can be directly called, omitting the embedding layer calculation process of the target word element that appears frequently, thereby saving computing resources and improving computing efficiency.

[0084] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0085] On the other hand, an embodiment of the present invention further provides a data processing device, such as Figure 7 As shown, the device includes:

[0086] A frequency information determination module 701 is configured to determine a first word-gram set and frequency information of each first word-gram in the first word-gram set based on word segmentation processing of the input first text segment;

[0087] A target word-gram determining module 702 is configured to take the first word-gram in the first word-gram set whose frequency information meets a preset condition as a target word-gram;

[0088] The embedding result storage module 703 is used to obtain the embedding result corresponding to the target word and store the embedding result corresponding to the target word in a preset storage area;

[0089] The mapping relationship establishing module 704 is used to establish a mapping relationship between the target word and the storage address of the corresponding embedding result, and the mapping relationship is used to search and call the embedding result corresponding to the target word.

[0090] In an optional implementation, the frequency information determination module 701 includes:

[0091] a word segmentation processing unit, configured to perform word segmentation processing on the input first text segment, determine word units contained in the first text segment, and update the first word unit set;

[0092] a frequency information to be updated determining unit, configured to count the number of occurrences of a first word contained in the first text segment in the first text segment to obtain the frequency information to be updated;

[0093] The frequency information updating unit is configured to obtain frequency information of a first word contained in the first text segment, and update the frequency information based on the frequency information to be updated.

[0094] In an optional implementation manner, the preset conditions include:

[0095] Frequency information reaches the preset threshold;

[0096] And / or, a ratio of the frequency information to the number of first word-grams in the first word-gram set reaches a preset ratio.

[0097] In an optional embodiment, the preset storage area includes at least one storage area, and the device further includes:

[0098] a storage quantity determination module, configured to determine the number of storage areas and the number of target word units corresponding to the storage areas based on a storage setting for the target word units, wherein the storage setting for the target word units includes setting a preset number of target word units and setting grouped storage of the target word units;

[0099] The storage size determination module is used to determine the size of the storage area based on the number of target word units corresponding to the storage area and the storage space requirements corresponding to the target word units.

[0100] In an optional embodiment, the device further comprises:

[0101] The frequency information determination module is further configured to determine, within each data evaluation period, a second word-gram set and frequency information of each second word-gram in the second word-gram set based on word segmentation processing of the input second text segment if the resource utilization rate of the preset storage area reaches a preset utilization rate, wherein the second word-gram set is used to collect statistics on the second word-grams within a data evaluation period;

[0102] a target word-gram updating module, configured to update the target word-gram based on the frequency information of the second word-gram in the second word-gram set;

[0103] The embedding result updating module is used to update the embedding result stored in the preset storage area based on the embedding result corresponding to the updated target word.

[0104] In an optional embodiment, the target word unit updating module includes:

[0105] a sorting unit, configured to sort the second word-grams in the second word-gram set based on the frequency information to obtain an arrangement order of the second word-grams in the second word-gram set;

[0106] The word unit extraction unit is used to extract a preset number of second word units from the second word unit set according to the arrangement order as target word units.

[0107] In an optional implementation, the embedded result update module includes:

[0108] A word-unit comparison unit, configured to compare the embedding result stored in the preset storage area with the updated target word-unit, and determine the word-unit to be stored and the storage area to be replaced;

[0109] The embedding result replacement unit is used to clear the embedding results stored in the storage area to be replaced, and store the embedding results corresponding to the word units to be stored in the storage area to be replaced.

[0110] For the description of the features in the embodiments corresponding to the data processing device, reference can be made to the relevant description of the embodiments corresponding to the data processing method, and no further details will be given here.

[0111] On the other hand, an embodiment of the present invention further provides an electronic device, such as Figure 8 As shown, one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the electronic device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 A processor 10 is taken as an example.

[0112] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0113] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0114] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0115] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0116] The electronic device further includes a communication interface 30 for the electronic device to communicate with other devices or a communication network.

[0117] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data processing method embodiments when run.

[0118] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0119] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above data processing method embodiments are implemented.

[0120] An embodiment of the present invention further provides another computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned data processing method embodiments are implemented.

[0121] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0122] The above is a detailed introduction to a data processing method provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core ideas of the present invention. It should be pointed out that, for those skilled in the art, without departing from the principles of the present invention, several improvements and modifications may be made to the present invention, and such improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A data processing method, characterized in that: include: Determining a first word-gram set and frequency information of each first word-gram in the first word-gram set based on word segmentation processing of the input first text segment; taking the first word-gram in the first word-gram set whose frequency information meets a preset condition as a target word-gram; The preset conditions include: the frequency information reaches a preset threshold, and / or the ratio of the frequency information to the number of the first word-grams in the first word-gram set reaches a preset ratio; Obtaining the embedding result corresponding to the target word, and storing the embedding result corresponding to the target word in a preset storage area; the embedding result corresponding to the target word is the calculation result output by the embedding layer after the target word is calculated by the embedding layer; A mapping relationship is established between the target word and the storage address of the corresponding embedding result, and the mapping relationship is used to search and call the embedding result corresponding to the target word.

2. The data processing method according to claim 1, wherein: The determining of the first word-gram set and the frequency information of each first word-gram in the first word-gram set based on the word segmentation processing of the input first text segment includes: Performing word segmentation processing on a first input text segment to determine a first word element contained in the first text segment to update the first word element set; Counting the number of occurrences of the first word contained in the first text segment in the first text segment to obtain frequency information to be updated; Frequency information of a first word contained in the first text segment is obtained, and the frequency information is updated based on the frequency information to be updated.

3. The data processing method according to claim 1, wherein: The preset storage area includes at least one storage area, and the method further includes: Determining the number of the storage areas and the number of target word units corresponding to the storage areas based on a storage setting for the target word units, wherein the storage setting for the target word units includes setting a preset number of the target word units and setting grouped storage of the target word units; The size of the storage area is determined based on the number of target word units corresponding to the storage area and the storage space requirements corresponding to the target word units.

4. The data processing method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the resource utilization rate of the preset storage area reaches a preset utilization rate, then within each data evaluation period, based on word segmentation processing of the input second text segment, determining a second word-gram set and frequency information of each second word-gram in the second word-gram set, wherein the second word-gram set is used to collect statistics on the second word-grams within one data evaluation period; updating the target word based on the frequency information of the second word in the second word set; Based on the updated embedding result corresponding to the target word, the embedding result stored in the preset storage area is updated.

5. The data processing method according to claim 4, characterized in that: The updating of the target word-gram based on the frequency information of the second word-gram in the second word-gram set includes: sorting the second word-grams in the second word-gram set based on the frequency information to obtain an arrangement order of the second word-grams in the second word-gram set; A preset number of second word-grams are extracted from the second word-gram set according to the arrangement order as the target word-grams.

6. The data processing method according to claim 4, characterized in that: The updating of the embedding result stored in the preset storage area based on the updated embedding result corresponding to the target word includes: Comparing the embedding result stored in the preset storage area with the updated target word-unit to determine the word-unit to be stored and the storage area to be replaced; The embedding results stored in the storage area to be replaced are cleared, and the embedding results corresponding to the word-unit to be stored are stored in the storage area to be replaced.

7. A data processing device, characterized in that: include: a frequency information determination module, configured to determine a first word-gram set and frequency information of each first word-gram in the first word-gram set based on word segmentation processing of the input first text segment; a target word-gram determining module, configured to take the first word-gram in the first word-gram set whose frequency information meets a preset condition as a target word-gram; The preset conditions include: the frequency information reaches a preset threshold, and / or the ratio of the frequency information to the number of the first word-grams in the first word-gram set reaches a preset ratio; An embedding result storage module is used to obtain the embedding result corresponding to the target word and store the embedding result corresponding to the target word in a preset storage area; the embedding result corresponding to the target word is the calculation result output by the embedding layer after the target word is calculated by the embedding layer; A mapping relationship establishment module is used to establish a mapping relationship between the target word and the storage address of the corresponding embedding result, and the mapping relationship is used to search and call the embedding result corresponding to the target word.

8. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data processing method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Lookup table loop language model

    CN117043859A