Data processing method and device, equipment and storage medium
By performing word segmentation processing and frequency statistics on text fragments of the large model, high-frequency word elements are selected and their embedded results are stored, the problems of large model's high consumption and low efficiency are solved, and the saving of computing resources and efficiency are achieved.
Patent Information
- Application Number
- CN202510797913.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-16
AI Technical Summary
In the calculation process, large models have problems such as high computing resources and low computing efficiency, especially the repeated embedding layer calculations during the inference process, resulting in waste of resources.
By performing word segmentation processing on the input text fragments, count the word element frequency, filter out the high-frequency word element, and store its embedding results in the preset storage area, establish a mapping relationship, and directly call the embedding results of the high-frequency word element, omitting the repeated embedding layer calculation.
Save computing resources and improve computing efficiency, especially on devices with limited computing power but large memory, significantly improve the inference speed of large models.
Smart Images

Figure CN120336510A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large models, and particularly to a data processing method, apparatus, device, and storage medium. Background Art
[0002] As a machine learning model with a large number of parameters and a complex computational structure, a large model can handle complex tasks and data, and has wide applications in various fields. However, the application of a large model depends on powerful computing capabilities. It consumes a large amount of computing resources during the training and inference processes, and the larger the scale of the large model, the more computing resources are required and the longer the computing time is. As a result, the large model has problems of high computing resource consumption and low computing efficiency. Summary of the Invention
[0003] The present invention provides a data processing method, apparatus, device, and storage medium to at least solve the problems of high computing resource consumption and low computing efficiency existing in large models in related technologies.
[0004] The present invention provides a data processing method, including: Based on the word segmentation processing of the input first text segment, determining a first set of word elements and the frequency information of each first word element in the first set of word elements; Taking the first word elements in the first set of word elements whose frequency information meets a preset condition as target word elements; Obtaining the embedding result corresponding to the target word element and storing the embedding result corresponding to the target word element in a preset storage area; Establishing a mapping relationship between the target word element and the storage address of the corresponding embedding result, where the mapping relationship is used to search for and call the embedding result corresponding to the target word element.
[0005] The present invention also provides a data processing apparatus, including: A frequency information determination module, configured to determine a first set of word elements and the frequency information of each first word element in the first set of word elements based on the word segmentation processing of the input first text segment; A target word element determination module, configured to take the first word elements in the first set of word elements whose frequency information meets a preset condition as target word elements; An embedding result storage module, configured to obtain the embedding result corresponding to the target word element and store the embedding result corresponding to the target word element in a preset storage area; A mapping relationship establishment module, configured to establish a mapping relationship between the target word element and the storage address of the corresponding embedding result, where the mapping relationship is used to search for and call the embedding result corresponding to the target word element.
[0006] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above data processing methods when executing the computer program.
[0007] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above data processing methods when executed by a processor.
[0008] The present invention also provides a computer program product including a computer program, which implements the steps of any of the above data processing methods when executed by a processor.
[0009] Through the present invention, the input first text segment is tokenized, the frequency information of the first tokens obtained by tokenization is counted, and the first tokens whose frequency information meets the preset conditions are used as target tokens, so as to screen out the tokens that appear frequently according to the statistics of the frequency information; at the same time, the embedding results corresponding to the target tokens are stored in a preset storage area, and a mapping relationship is established between the target tokens and the storage addresses of the corresponding embedding results, so that the embedding results of the target tokens that appear frequently can be directly called when reasoning about the text segment later, omitting the calculation process of the embedding layer of the target tokens that appear frequently, thereby saving computing resources and improving computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0011] Figure 1 is a flowchart of the data processing method according to the embodiment of the present invention; Figure 2 is a schematic diagram of a preset storage area in the data processing method according to the embodiment of the present invention; Figure 3 is a flowchart of another data processing method according to the embodiment of the present invention; Figure 4 is a flowchart of yet another data processing method according to the embodiment of the present invention; Figure 5 is a schematic diagram of a specific embodiment of the data processing method according to the embodiment of the present invention; Figure 6 is a schematic diagram of another specific embodiment of the data processing method according to the embodiment of the present invention; Figure 7 is a block diagram of the structure of the data processing device according to the embodiment of the present invention; Figure 8 It is a schematic diagram of the hardware structure of the electronic device according to an embodiment of the present invention. Detailed implementation manners
[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0013] It should be noted that in the description of the present invention, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0014] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0015] As a machine learning model with a large number of parameters and a complex computing structure, the large model can handle complex tasks and data and has wide applications in various fields. However, the application of the large model depends on powerful computing capabilities. It consumes a large amount of computing resources during the training and inference processes, and the larger the scale of the large model, the more computing resources are required and the longer the computing time is, resulting in problems of high computing resource consumption and low computing efficiency in the large model.
[0016] Specifically, the inference process of the large model is actually a process of predicting the next token based on the previously input or generated tokens. Therefore, each prediction requires the completion of the forward calculation of the current last token. The currently generated token is first processed by a tokenizer into a token number of integer type, passes through an embedding layer, and changes from an integer to a tensor; then, each layer of the transformer structure calculates the tensor, where each layer undergoes attention and multi-layer perceptron (mlp) processing, and finally the calculation result is output. In fact, before the calculation of the transformer, the embedding results obtained by the same token through the embedding operation are the same. However, in the inference process of the large model, when a token appears again, it will still undergo repeated embedding layer calculations during the inference process, resulting in a waste of computing resources.
[0017] In the related art, there is a solution that uses the Multilingual BERT model to prune the embedding layer of low-frequency tokens. This solution tokenizes the target corpus during the inference process through the vocabulary of the BERT model, counts the token frequencies, and prunes the embedding layer of tokens with extremely low frequencies, thereby reducing the number of parameters in the embedding layer. However, such a method only reduces the number of parameters of the model and reduces the memory usage consumption, but cannot reduce the computational operations during the inference process and cannot improve the computational efficiency of the large model inference process.
[0018] Based on this, the present invention provides a data processing method. Based on the tokenization processing of the input first text segment, a first token set and the frequency information of each first token in the first token set are determined; the first tokens in the first token set whose frequency information meets the preset conditions are used as target tokens; the embedding results corresponding to the target tokens are obtained, and the embedding results corresponding to the target tokens are stored in a preset storage area; a mapping relationship is established between the target tokens and the storage addresses of the corresponding embedding results, and the mapping relationship is used to search for and call the embedding results corresponding to the target tokens. Thus, the input first text segment is tokenized, the frequency information of the first tokens obtained by tokenization is counted, and the first tokens whose frequency information meets the preset conditions are used as target tokens, so as to screen out the tokens that appear frequently according to the statistics of the frequency information; at the same time, the embedding results corresponding to the target tokens are stored in a preset storage area, and a mapping relationship is established between the target tokens and the storage addresses of the corresponding embedding results, so that when the text segment is inferred later, the embedding results of the target tokens that appear frequently can be directly called, omitting the embedding layer calculation process of the target tokens that appear frequently, thereby saving computing resources and improving computing efficiency.
[0019] The data processing method provided by the embodiments of the present invention can be applied not only to large models, but also to other models with embedding calculations. Especially for devices with limited computing power but large memory, it can well improve the computing efficiency. At the same time, if the applied model structure is a model structure that does not require rotational position encoding calculation, the final calculation results of the target tokens can be directly stored in the preset storage area, thereby directly omitting the model inference calculation process and further improving the computing efficiency.
[0020] In this embodiment, a data processing method is provided, which can be used for various models with embedding calculations. Figure 1 It is a schematic flowchart of the data processing method of the embodiments of the present invention. As Figure 1 shown, the process includes the following steps: Step S101, based on the tokenization processing of the input first text segment, determine a first token set and the frequency information of each first token in the first token set.
[0021] In an embodiment of the present invention, tokenization is performed on the first text segment input into the model. Any tokenization method can be used for tokenization, and no specific limitation is imposed here. The tokens obtained by tokenizing the first text segment are used as the first tokens to form a first token set. At the same time, the occurrence times of each first token in the first text segment are counted to obtain the frequency information of each first token.
[0022] Step S102: Use the first tokens in the first token set whose frequency information meets the preset conditions as target tokens.
[0023] In an embodiment of the present invention, the target tokens are tokens that frequently appear during the inference process. The embedding results corresponding to the target tokens will be stored in a preset storage area for easy searching and calling. Here, the embedding result corresponding to the target token refers to the calculation result output after the target token passes through the calculation of the embedding layer.
[0024] In an alternative embodiment, the preset conditions include: the frequency information reaches a preset threshold; and / or, the ratio of the frequency information to the number of first tokens in the first token set reaches a preset ratio. Here, the preset threshold can be set by the user or can be a fixed value set in advance.
[0025] Step S103: Obtain the embedding results corresponding to the target tokens and store the embedding results corresponding to the target tokens in a preset storage area.
[0026] In an embodiment of the present invention, after the target tokens pass through the calculation of the embedding layer, the calculation results output by the embedding layer are obtained to obtain the embedding results corresponding to the target tokens, and the embedding results corresponding to the target tokens are stored in a preset storage area. Here, the preset storage area is a storage area that is pre-divided in the memory and is used to store the embedding results.
[0027] In an alternative embodiment, the preset storage area includes at least one storage area. The at least one storage area can be distributed dispersedly in the memory, so that storage areas can be allocated at various positions in the memory, making full use of the memory fragments. Figure 2 is a schematic diagram of the preset storage area in the data processing method of the embodiment of the present invention. As Figure 2 shown, the preset storage area can be divided into multiple storage areas and is distributed dispersedly in the unused memory according to the size of the memory fragments. The No. 1 storage area is located in one memory fragment, and the No. 2 and No. 3 memory areas are located in another memory fragment, so as to make full use of the memory fragments.
[0028] Correspondingly, the preset storage area can be divided and allocated in the following manner: First, based on the storage settings for the target tokens, determine the number of storage areas and the number of target tokens corresponding to each storage area. Among them, the storage settings for the target tokens include the setting of the preset number of target tokens and the setting of grouped storage of target tokens. The setting of the preset number of target tokens limits the number of high-frequency tokens to be stored, and the setting of grouped storage of target tokens limits the number of embedding results corresponding to the target tokens stored in one storage area. Dividing the latter by the former can obtain the number of storage areas. For example, the preset number of target tokens can be set to 40, and 4 target tokens can be set as a group, then 10 storage areas need to be divided. Then, based on the number of target tokens corresponding to the storage area and the storage space requirements corresponding to the target tokens, determine the size of the storage area. For example, if 4 target tokens are set as a group and each F32 data format occupies 4 bytes, then the size of one storage area is 4 * number of embedding layers * 4 Bytes.
[0029] Step S104, establish a mapping relationship between the target token and the storage address of the corresponding embedding result.
[0030] In the embodiment of the present invention, the mapping relationship between the target token and the storage address of the corresponding embedding result is used to search for and call the embedding result corresponding to the target token. In the subsequent reasoning process of the newly input text segment, it is possible to first determine whether the token obtained by word segmentation is a target token by searching the mapping relationship. If it is a target token, the storage address of the corresponding embedding result can be determined through the mapping relationship, and the data at the corresponding storage address can be read from the memory to obtain the embedding result, thus omitting the calculation process of the embedding layer and improving the calculation efficiency.
[0031] In an optional implementation manner, a storage list can be pre-constructed, and the storage addresses corresponding to the preset storage areas and the target tokens corresponding to the embedding results stored in the storage addresses are recorded in the storage list. When storing the embedding result corresponding to the target token into the preset storage area, in the storage list, fill in the target token at the position corresponding to the storage address where the embedding result corresponding to the target token is stored, so as to establish the mapping relationship between the target token and the storage address of the corresponding embedding result. In the subsequent reasoning process, the embedding result corresponding to the target token can be quickly located by directly searching the storage list.
[0032] The data processing method provided by the embodiments of the present invention determines a first token set and the frequency information of each first token in the first token set based on word segmentation processing of the input first text segment; uses the first tokens in the first token set whose frequency information meets a preset condition as target tokens; obtains the embedding results corresponding to the target tokens, and stores the embedding results corresponding to the target tokens in a preset storage area; establishes a mapping relationship between the target tokens and the storage addresses of the corresponding embedding results, and the mapping relationship is used to search for and call the embedding results corresponding to the target tokens. Thus, word segmentation processing is performed on the input first text segment, the frequency information of the first tokens obtained by word segmentation is counted, and the first tokens whose frequency information meets the preset condition are used as target tokens, so as to screen out the tokens that appear frequently according to the statistics of the frequency information; at the same time, the embedding results corresponding to the target tokens are stored in a preset storage area, and a mapping relationship is established between the target tokens and the storage addresses of the corresponding embedding results, so that the embedding results of the target tokens that appear frequently can be directly called when performing inference on the text segment subsequently, omitting the calculation process of the embedding layer of the target tokens that appear frequently, thereby saving computing resources and improving computing efficiency.
[0033] In this embodiment, a data processing method is provided, which can be used in various models with embedding calculations. Figure 3 It is a flowchart of another data processing method of the embodiments of the present invention, as Figure 3 shown, and the process includes the following steps: Step 301: Based on word segmentation processing of the input first text segment, determine a first token set and the frequency information of each first token in the first token set.
[0034] In the embodiments of the present invention, when the preset storage area for storing embedding results is not full or the resource utilization rate does not reach the preset utilization rate, the text segment input to the model is used as the first text segment, and for the first text segment, the first tokens obtained by its word segmentation processing will be cumulatively counted. Correspondingly, step S301 includes: Step S3011: Perform word segmentation processing on the input first text segment to determine the first tokens included in the first text segment, so as to update the first token set.
[0035] In the embodiments of the present invention, when the preset storage area for storing embedding results is not full or the resource utilization rate does not reach the preset utilization rate, whenever a new first text segment is input to the model, word segmentation processing is performed on the input first text segment to obtain the first tokens included in the first text segment, and compare them with the first tokens included in the first token set to determine the first tokens that are not stored in the first token set among the first tokens included in the first text segment, that is, the first tokens that appear for the first time, and update them to the first token set.
[0036] Step S3012: Count the number of occurrences of the first token included in the first text segment to obtain the frequency information to be updated.
[0037] In the embodiments of the present invention, the number of occurrences of each first token included in the input first text segment is counted separately, and the number of occurrences of the first token in the first text segment is used as the frequency information to be updated corresponding to the first token.
[0038] Step S3013: Obtain the frequency information of the first tokens included in the first text segment, and update the frequency information based on the frequency information to be updated.
[0039] In the embodiments of the present invention, the frequency information of the first tokens included in the first text segment is obtained. This frequency information is the statistics of the total number of occurrences of the first tokens in all input first text segments at the historical moment; on the basis of this frequency information, the frequency information to be updated is superimposed to update the frequency information.
[0040] Step S302: Use the first tokens in the first token set whose frequency information meets the preset conditions as target tokens. For details, please refer to Figure 1 Step S102 of the illustrated embodiment, which will not be elaborated here.
[0041] Step S303: Obtain the embedding results corresponding to the target tokens, and store the embedding results corresponding to the target tokens in a preset storage area. For details, please refer to Figure 1 Step S103 of the illustrated embodiment, which will not be elaborated here.
[0042] Step S304: Establish a mapping relationship between the target tokens and the storage addresses of the corresponding embedding results. For details, please refer to Figure 1 Step S104 of the illustrated embodiment, which will not be elaborated here.
[0043] The data processing method provided by the embodiments of the present invention performs word segmentation on the input first text segment, determines the first tokens included in the first text segment to update the first token set, counts the number of occurrences of the first tokens included in the first text segment in the first text segment to obtain the frequency information to be updated, obtains the frequency information of the first tokens included in the first text segment, and updates the frequency information based on the frequency information to be updated, so as to continuously update the first token set and the frequency information of each first token based on the newly input first text segment, and ensure the accuracy and reliability of the target tokens determined based on the frequency information.
[0044] In this embodiment, a data processing method is provided, which can be used in various models with embedding calculations. Figure 4 It is a flowchart of another data processing method of the embodiments of the present invention, asFigure 4 As shown in the figure, the process includes the following steps: Step S401: Based on the word segmentation processing of the input first text fragment, determine the first token set and the frequency information of each first token in the first token set. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.
[0045] Step S402: Use the first tokens in the first token set whose frequency information meets the preset conditions as target tokens. For details, please refer to Figure 1 Step S102 of the embodiment shown, which will not be elaborated here.
[0046] Step S403: Obtain the embedding result corresponding to the target token and store the embedding result corresponding to the target token in the preset storage area. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.
[0047] Step S404: Establish a mapping relationship between the target token and the storage address of the corresponding embedding result. For details, please refer to Figure 1 Step S104 of the embodiment shown, which will not be elaborated here.
[0048] Step S405: If the resource utilization rate of the preset storage area reaches the preset utilization rate, then in each data evaluation period, based on the word segmentation processing of the input second text fragment, determine the second token set and the frequency information of each second token in the second token set.
[0049] In the embodiment of the present invention, if the resource utilization rate of the preset storage area reaches the preset utilization rate, it indicates that the preset storage area can no longer store more embedding results corresponding to the target tokens. At this time, an optimization strategy of storage competition is adopted to ensure that the embedding results corresponding to the tokens that appear most frequently are always stored in the preset storage area, thereby maximizing the utilization efficiency of the storage resources.
[0050] In the embodiment of the present invention, in the optimization strategy of storage competition, in each data evaluation period, based on the word segmentation processing of the input second text fragment, determine the second token set and the frequency information of each second token in the second token set. Among them, after enabling the optimization strategy of storage competition, that is, when the resource utilization rate of the preset storage area reaches the preset utilization rate, the text fragments input in each data evaluation period are respectively used as the second text fragment, that is, the text fragments are separately calculated and statistically analyzed in each data evaluation period; correspondingly, the second token set is used to count the second tokens in a data evaluation period.
[0051] In the embodiments of the present invention, during each data evaluation period, the input second text segment is tokenized to determine the second token set and the frequency information of each second token in the second token set. Specifically, reference can be made to Figure 3 Step S301 in the illustrated embodiment, which will not be elaborated here.
[0052] In an alternative embodiment, the data evaluation period can be determined based on the number of rounds of the inference task, that is, within one data evaluation period, the model will complete a preset number of rounds of inference tasks.
[0053] Step S406, update the target token based on the frequency information of the second tokens in the second token set.
[0054] In the embodiments of the present invention, based on the frequency information of the second tokens in the second token set, the second tokens that frequently appear in the second token set are determined and used as the target tokens to update the target tokens.
[0055] In an alternative embodiment, the following method can be used to update the target token based on the frequency information of the second tokens in the second token set: Sort the second tokens in the second token set based on the frequency information to obtain the arrangement order of the second tokens in the second token set; among them, the second tokens in the second token set are sorted according to the magnitude of the frequency information, and the second tokens are arranged in the descending or ascending order of the frequency information. According to the arrangement order, a preset number of second tokens are extracted from the second token set as the target tokens; among them, if the second tokens are arranged in the descending order of the frequency information, a preset number of second tokens are extracted from the front to the back according to the arrangement order as the target tokens, and if the second tokens are arranged in the ascending order of the frequency information, a preset number of second tokens are extracted from the back to the front according to the arrangement order as the target tokens, so as to extract the preset number of second tokens with the highest occurrence frequency in the second token set and ensure that the embedding results corresponding to the most frequently occurring tokens are stored.
[0056] In an alternative embodiment, second tokens whose frequency information meets the preset conditions can also be selected from the second token set as the target tokens. If the number of target tokens selected from the second token set does not reach the upper limit of the number of target tokens, that is, the preset number, then the target tokens corresponding to the embedding results stored in the preset storage area are screened according to the frequency information, and the target tokens with lower frequency information are screened out. The number of screened tokens is the same as the number of target tokens selected from the second token set, and the screened target tokens are cleared to release storage space for the target tokens selected from the second token set.
[0057] Step S407: Update the embedding results stored in the preset storage area based on the embedding results corresponding to the updated target tokens.
[0058] In the embodiments of the present invention, considering that there may be partial overlap between the target tokens corresponding to the embedding results already stored in the preset storage area and the updated target tokens, the embedding results corresponding to the overlapping target tokens in the preset storage area are retained, and the storage space of the embedding results corresponding to the non-overlapping target tokens in the preset storage area is released.
[0059] Specifically, first, compare the embedding results stored in the preset storage area with the updated target tokens to determine the tokens to be stored and the storage areas to be replaced; where the tokens to be stored refer to the tokens in the target tokens whose corresponding embedding results are not stored in the preset storage area, and the storage areas to be replaced refer to the storage areas in the preset storage area where the tokens corresponding to the stored embedding results are no longer the target tokens, and the embedding results stored in the areas to be stored need to be cleared to release the storage space. Then, clear the embedding results stored in the storage areas to be replaced, and store the embedding results corresponding to the tokens to be stored in the storage areas to be replaced.
[0060] The data processing method provided by the embodiments of the present invention, when the resource utilization rate of the preset storage area reaches the preset utilization rate, within each data evaluation period, based on the word segmentation processing of the input second text segment, determine the second token set and the frequency information of each second token in the second token set, and update the target tokens based on the frequency information of the second tokens in the second token set, so as to determine the second token set and the frequency information of each second token in the second token set according to the data evaluation period, so as to be able to dynamically adjust the frequently occurring tokens stored in real time according to the inference task situation of the model; at the same time, update the embedding results stored in the preset storage area based on the embedding results corresponding to the updated target tokens, so as to ensure that the preset storage area always stores the embedding results corresponding to the most frequently occurring tokens, so that the embedding results stored in the preset storage area can match the inference task of the actual model, ensure that the embedding results stored in the preset storage area can play the greatest role, and improve the computing efficiency of the model.
[0061] As a specific implementation manner, Figure 5 is a schematic diagram of a specific embodiment of the data processing method of the embodiments of the present invention, Figure 5 shows the processing process of tokens when the preset storage area is not fully stored, that is, when the resource utilization rate of the preset storage area does not reach the preset utilization rate. As Figure 5As shown, on the processor (CPU) side, it is determined whether the token obtained by segmenting the text fragment is the target token. On the graphics processing unit (GPU) side, the token is processed into a token number of integer type by the tokenizer component, and then through the embedding layer, the token is changed from an integer to a tensor to obtain the embedding result for the subsequent model inference process.
[0062] Among them, first, on the processor (CPU) side, it is determined whether the token is already stored, that is, whether the token has been determined as the target token. If so, on the graphics processing unit (GPU) side, the embedding result is directly obtained from the preset storage area for the subsequent model inference process; if not, the count of this token is incremented by 1 for frequency information accumulation and update, and after the update, based on the frequency information, it is determined whether this token meets the storage condition and can be determined as the target token. If the storage condition is met, a storage address is allocated for it from the preset storage area, and then on the graphics processing unit (GPU) side, the token is processed into a token number of integer type by the tokenizer component, and through the embedding layer, the token is changed from an integer to a tensor to obtain the embedding result, and the embedding result is stored at the allocated storage address, and the subsequent model inference process is carried out; if the storage condition is not met, directly on the graphics processing unit (GPU) side, the token is processed into a token number of integer type by the tokenizer component, and through the embedding layer, the token is changed from an integer to a tensor to obtain the embedding result and carry out the subsequent model inference process.
[0063] As a specific implementation manner, Figure 6 is a schematic diagram of another specific embodiment of the data processing method of the embodiment of the present invention. Figure 6 shows the processing process of tokens in the case where the preset storage area is full, that is, the resource utilization rate of the preset storage area reaches the preset utilization rate. As Figure 6 shown, within a data evaluation cycle, the specified number of rounds of inference tasks are completed, the occurrence frequencies of the tokens that appear are counted and sorted, and the high-frequency tokens, that is, the target tokens, are determined according to the sorting order. Then, it is judged whether the token corresponding to the embedding result stored in the preset storage area is still the target token. If so, the corresponding embedding result stored in the preset storage area is retained. If not, the corresponding embedding result stored in the preset storage area is cleared to release the preset storage area, and the embedding result corresponding to the un-stored target token is stored in the preset storage area, thereby completing the determination of the target token and the storage of the embedding result within a data evaluation cycle.
[0064] The data processing method provided by the embodiments of the present invention is based on word segmentation processing of the input first text segment to determine a first token set and the frequency information of each first token in the first token set; the first tokens in the first token set whose frequency information meets the preset conditions are used as target tokens; the embedding results corresponding to the target tokens are obtained and stored in a preset storage area; a mapping relationship is established between the target tokens and the storage addresses of the corresponding embedding results, and the mapping relationship is used to search for and call the embedding results corresponding to the target tokens. Thus, word segmentation processing is performed on the input first text segment, the frequency information of the first tokens obtained by word segmentation is counted, and the first tokens whose frequency information meets the preset conditions are used as target tokens, so as to screen out the tokens that appear frequently according to the statistics of the frequency information; at the same time, the embedding results corresponding to the target tokens are stored in a preset storage area, and a mapping relationship is established between the target tokens and the storage addresses of the corresponding embedding results, so that when inferring the text segment subsequently, the embedding results of the target tokens that appear frequently can be directly called, omitting the calculation process of the embedding layer of the target tokens that appear frequently, thereby saving computing resources and improving computing efficiency.
[0065] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0066] On the other hand, the embodiments of the present invention also provide a data processing device, as Figure 7 shown, the device includes: A frequency information determination module 701, configured to determine a first token set and the frequency information of each first token in the first token set based on word segmentation processing of the input first text segment; A target token determination module 702, configured to use the first tokens in the first token set whose frequency information meets the preset conditions as target tokens; An embedding result storage module 703, configured to obtain the embedding results corresponding to the target tokens and store the embedding results corresponding to the target tokens in a preset storage area; A mapping relationship establishment module 704, configured to establish a mapping relationship between the target tokens and the storage addresses of the corresponding embedding results, and the mapping relationship is used to search for and call the embedding results corresponding to the target tokens.
[0067] In an optional implementation manner, the frequency information determination module 701 includes: A word segmentation processing unit, configured to perform word segmentation processing on the input first text segment to determine the tokens included in the first text segment, so as to update the first token set; A to-be-updated frequency information determining unit, configured to count the number of occurrences of a first token included in a first text segment in the first text segment, to obtain to-be-updated frequency information; A frequency information updating unit, configured to obtain the frequency information of the first token included in the first text segment, and update the frequency information based on the to-be-updated frequency information.
[0068] In an alternative embodiment, the preset conditions include: The frequency information reaches a preset threshold; And / or, the ratio of the frequency information to the number of first tokens in the first token set reaches a preset ratio.
[0069] In an alternative embodiment, the preset storage area includes at least one storage area, and the apparatus further includes: A storage quantity determining module, configured to determine the number of storage areas and the number of target tokens corresponding to the storage areas based on the storage settings for the target tokens. The storage settings for the target tokens include the setting of a preset quantity for the target tokens and the setting of grouped storage for the target tokens; A storage size determining module, configured to determine the size of the storage area based on the number of target tokens corresponding to the storage area and the storage space requirements corresponding to the target tokens.
[0070] In an alternative embodiment, the apparatus further includes: A frequency information determining module, further configured to, if the resource utilization rate of the preset storage area reaches a preset utilization rate, within each data evaluation period, based on the word segmentation processing of the input second text segment, determine a second token set and the frequency information of each second token in the second token set, where the second token set is used to count the second tokens within one data evaluation period; A target token updating module, configured to update the target tokens based on the frequency information of the second tokens in the second token set; An embedding result updating module, configured to update the embedding results stored in the preset storage area based on the embedding results corresponding to the updated target tokens.
[0071] In an alternative embodiment, the target token updating module includes: A sorting unit, configured to sort the second tokens in the second token set based on the frequency information, to obtain the arrangement order of the second tokens in the second token set; A token extraction unit, configured to extract a preset number of second tokens from the second token set according to the arrangement order, as the target tokens.
[0072] In an alternative embodiment, the embedding result updating module includes: A token comparison unit, configured to compare the embedding result stored in a preset storage area with an updated target token to determine a token to be stored and a storage area to be replaced; An embedding result replacement unit, configured to clear the embedding result stored in the storage area to be replaced and store the embedding result corresponding to the token to be stored in the storage area to be replaced.
[0073] For the description of the features in the corresponding embodiments of the data processing device, reference may be made to the relevant description in the corresponding embodiments of the data processing method, which will not be elaborated here one by one.
[0074] On the other hand, an embodiment of the present invention further provides an electronic device, as Figure 8 shown, including one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 8 Taking one processor 10 as an example.
[0075] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.
[0076] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0077] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0078] The memory 20 may include a volatile memory, for example, a random access memory; the memory may also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.
[0079] The electronic device further includes a communication interface 30 for the electronic device to communicate with other devices or communication networks.
[0080] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any one of the above data processing method embodiments when running.
[0081] In an exemplary embodiment, the above computer-readable storage medium may include but is not limited to: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0082] An embodiment of the present invention further provides a computer program product, where the above computer program product includes a computer program, and the computer program implements the steps in any one of the above data processing method embodiments when executed by a processor.
[0083] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and the computer program implements the steps in any one of the above data processing method embodiments when executed by a processor.
[0084] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0085] The above has introduced in detail a data processing method provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A data processing method, characterized in that, Including: Based on the word segmentation of the input first text segment, determine the first token set and the frequency information of each first token in the first token set; Take the first tokens in the first token set whose frequency information meets the preset conditions as target tokens; Obtain the embedding results corresponding to the target tokens and store the embedding results corresponding to the target tokens in a preset storage area; Establish a mapping relationship between the target tokens and the storage addresses of the corresponding embedding results, and the mapping relationship is used to search for and call the embedding results corresponding to the target tokens.
2. The data processing method according to claim 1, wherein The step of "Based on the word segmentation of the input first text segment, determine the first token set and the frequency information of each first token in the first token set" includes: Perform word segmentation on the input first text segment to determine the first tokens included in the first text segment, so as to update the first token set; Count the number of occurrences of the first tokens included in the first text segment in the first text segment to obtain the frequency information to be updated; Obtain the frequency information of the first tokens included in the first text segment, and update the frequency information based on the frequency information to be updated.
3. The data processing method according to claim 1, characterized in that The preset conditions include: The frequency information reaches a preset threshold; And / or, the ratio of the frequency information to the number of the first tokens in the first token set reaches a preset ratio.
4. The data processing method according to claim 1, wherein The preset storage area includes at least one storage area, and the method further includes: Based on the storage settings for the target tokens, determine the number of the storage areas and the number of target tokens corresponding to the storage areas. The storage settings for the target tokens include the setting of the preset number of the target tokens and the setting of storing the target tokens in groups; Determine the size of the storage area based on the number of target tokens corresponding to the storage area and the storage space requirements corresponding to the target tokens.
5. The data processing method according to any one of claims 1-4, characterized in that The method further includes: If the resource utilization rate of the preset storage area reaches a preset utilization rate, then in each data evaluation period, based on the word segmentation of the input second text segment, determine the second token set and the frequency information of each second token in the second token set, where the second token set is used to count the second tokens in one data evaluation period; Update the target tokens based on the frequency information of the second tokens in the second token set; Update the embedding results stored in the preset storage area based on the embedding results corresponding to the updated target tokens.
6. The data processing method according to claim 5, characterized in that The step of "Update the target tokens based on the frequency information of the second tokens in the second token set" includes: Sort the second tokens in the second token set based on the frequency information to obtain the arrangement order of the second tokens in the second token set; Extract a preset number of second tokens from the second token set according to the arrangement order as the target tokens.
7. The data processing method according to claim 5, characterized in that, The step of "Update the embedding results stored in the preset storage area based on the embedding results corresponding to the updated target tokens" includes: Compare the embedding result stored in the preset storage area with the updated target token to determine the token to be stored and the storage area to be replaced; Clear the embedding result stored in the storage area to be replaced, and store the embedding result corresponding to the token to be stored in the storage area to be replaced.
8. A data processing device, characterized in that, Comprising: A frequency information determination module, configured to determine a first token set and the frequency information of each first token in the first token set based on word segmentation processing of an input first text segment; A target token determination module, configured to use the first token in the first token set whose frequency information meets a preset condition as the target token; An embedding result storage module, configured to obtain the embedding result corresponding to the target token and store the embedding result corresponding to the target token in a preset storage area; A mapping relationship establishment module, configured to establish a mapping relationship between the target token and the storage address of the corresponding embedding result, where the mapping relationship is used to search for and call the embedding result corresponding to the target token.
9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the data processing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the data processing method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Fine adjustment method and device of BERT model, equipment and storage medium
CN116933862A
Lookup table loop language model
CN117043859A
Model reasoning method and device, electronic equipment and storage medium
CN119129750A
Model fine tuning method, text processing method, medium, equipment and program product
CN119150862A
Language task processing method, system and device, storage medium and program product
CN120068846A
Cited By
Communication and calculation parallel method and device, equipment and medium
CN122111700A
Probiotic tolerance characteristic prediction method and device, electronic equipment and storage medium
CN122117042A