Information processing device

WO2025094972A1PCT designated stage expired Publication Date: 2025-05-08PREFERRED ELEMENTS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/038632
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively control the information embedded in text data generated by large-scale language models.

Method used

By using multiple processors and memory in the information processing device, the token generated by the generation model is selected, and the token sequence is controlled based on the embedding information, and the embedding information is embedded as a digital watermark.

Benefits of technology

The fine control of embedded information in the generated data of the generative model is realized, ensuring that the embedded information can be effectively embedded and read, thus solving the problem of uncontrolled information embedding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024038632_08052025_PF_FP_ABST
    Figure JP2024038632_08052025_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To determine where information is to be provided. [Solution] This information processing device comprises at least one memory and at least one processor. The at least one processor selects, on the basis of embedding information, a token to be included in a token sequence to be outputted as data from among tokens generated by a generative model, and embeds the embedding information in the data.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device

[0001] The present disclosure relates to an information processing device.

[0002] When generating text from a large-scale language model, there is a technology that embeds a digital watermark in the text. If the digital watermark is detected in the text, it can be determined that the text was generated from a large-scale language model.

[0003] J. Kirchenbauer, et.al., "A Watermark for Large Language Models", 25 April 2023, las modified 16 June 2023, ICML 2023, https: / / openreview.net / forum?id=aX8ig9X2a7

[0004] One non-limiting problem that embodiments of the present disclosure seek to solve is controlling the information that is embedded in the data generated by a generative model.

[0005] According to one embodiment, an information processing device includes one or more memories and one or more processors, which select tokens to be included in a token string output as data from among tokens generated by a generative model based on embedding information, and embed the embedded information in the data.

[0006] FIG. 1 is a diagram schematically illustrating an example of token generation according to an embodiment. FIG. 2 is a diagram schematically illustrating an example of a token sequence according to an embodiment. FIG. 3 is a diagram schematically illustrating an example of a token sequence according to an embodiment. FIG. 4 is a diagram schematically illustrating an example of embedded information acquisition according to an embodiment. FIG. 5 is a diagram schematically illustrating an application example as a non-limiting example according to an embodiment. FIG. 6 is a diagram illustrating an example of implementation of an information processing device according to an embodiment.

[0007] The problems to be solved by the embodiments of the present disclosure are not limited to the problems described above, and as further examples of some problems that are not limited to these, problems corresponding to the effects described in the embodiments can also be considered. In other words, a problem corresponding to at least one of the effects described in the description of the embodiments of the present disclosure can be considered to be a problem to be solved by the present disclosure.

[0008] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described with reference to the accompanying drawings, which are given by way of example only and are not intended to limit the scope of the present invention.

[0009] For example, in this disclosure, a case will be described in which text data such as characters and sentences, which are information, is used as a language model, but the embodiment of the present invention is not limited to this and can also be applied to data such as images, audio, video, sensor data, etc. These data can be similarly applied in the embodiment by being appropriately tokenized.

[0010] This disclosure describes, by way of non-limiting examples, information processing devices that embed information in data and information processing devices that read information from data.

[0011] The information processing device according to the present disclosure includes one or more memories and one or more processors. In addition, the information processing device may include a power supply unit, an interface, or other necessary components required for the operation of the information processing device. Hereinafter, the one or more memories may be simply referred to as memories, and the one or more processors may be simply referred to as processors.

[0012] At least one of the one or more memories may be external to the information processing device, and at least one of the one or more processors may be external to the information processing device. When at least one memory or at least one processor is external, the information processing device described as a non-limiting example in this disclosure can operate as at least a part of an information processing system.

[0013] (Embodiment relating to an embedding device)

[0014] As a non-limiting example, the information processing device generates data using a generative model of some kind of data and embeds information in the generated data. The generative model may be, for example, a large-scale language model, but is not limited thereto. The generative model according to the present disclosure may be, for example, a model that can divide the generated data into tokens.

[0015] The information processing device may also be configured to embed information into data using a model that modifies data. As described above, this model may be, for example, a model that can divide data to be modified into tokens. By using this model, for example, the information processing device may embed information into data generated using the above-described generative model after the data is generated.

[0016] When using the modification model, the "data to be generated" in the following description can be read as "data to be modified." For example, the information processing device can modify some of the tokens of data that can be divided into tokens, and embed the embedded information into the target data.

[0017] It should also be noted that in the present disclosure, embedding information in data and embedding information in a token sequence that constitutes the data may be the same process.

[0018] FIG. 1 is a diagram illustrating a process of generating tokens according to an embodiment. Tokens generated up to now are s0 to s t-1 The processor is denoted by the previous token s t-1 is the first token, and the second token s is based on this first token. t Generate.

[0019] The output data output by the processor is, for example, from s0 to the final token s N This is the data that is concatenated up to s. N may be a token indicating the end of the data. In this case, the output data is N tokens s0 to s N-1It is represented as a concatenated token.

[0020] The processor embeds one piece of information (hereinafter referred to as "embedded information") into these N tokens. The embedded information may be any information. The generation of data with this embedded information will be described in detail below.

[0021] The processor generates multiple candidates for the next token (token candidates) c0, c1, c2, ..., based on the past tokens (S100). This generation is not particularly limited, and the processor may generate them using a rule base or a generative model, for example. That is, the second token s t The candidates for are at least the first token s generated in the past. t-1 is generated based on

[0022] The processor generates a first token s in parallel with the generation of the token candidates, or before or after the generation of the token candidates. t-1 The processor generates a hash seed S from the first token s t-1 is converted into a hash value using a predefined hash seed, and based on this hash value, a hash seed S, which is a seed used to convert (hash) each token candidate into a hash value, can be calculated. The conversion from the hash value to the hash seed can be any reversible or irreversible conversion. This conversion may also be linear or nonlinear.

[0023] The processor calculates hash values ​​h0, h1, h2, ... for each of the multiple candidates c0, c1, c2, ... using the hash seed S (S104). That is, the processor hashes each of the multiple candidates.

[0024] The processor converts the plurality of second token candidates according to a predetermined rule to obtain corresponding values, as described above. That is, the predetermined rule may be, as a non-limiting example, a conversion that includes at least calculation of a hash value.

[0025] The calculation of this hash value is performed, for example and not by way of limitation, using a predetermined hash seed and the first token s t-1 This can include a transformation using

[0026] As a more specific example, but not limited to, the first token s t-1 A value related to the converted hash value can be used.

[0027] As a further non-limiting example, the predetermined rule may be a transformation that uses the acquired hash seed S to obtain a candidate hash value.

[0028] The above is merely an example, and does not exclude other methods. In another non-limiting example, the processor may use a predetermined rule to convert a plurality of candidates in a way that is unbiased or has a small bias with respect to some index.

[0029] The processor classifies the second token candidates c0, c1, c2, ... into groups based on the hash values ​​h0, h1, h2, ... (S106). Any method can be used as the classification rule. The processor classifies the second token candidates into groups based on any classification method, such as using a predetermined threshold for the hash values, using a method such as even / odd for the hash values, or classifying the hash values ​​by a predetermined method.

[0030] The group may be, for example, two groups, group A and group B, as shown in the figure. As an example, the processor may define 0 in group A and 1 in group B. That is, in a non-limiting example, the processor classifies candidates into a group indicating a value of 0 and a group indicating a value of 1 based on the conversion by S104. These "0" and "1" may correspond to bit values.

[0031] It is possible to include this classification in the predetermined rule along with the processes of S102 and S104 above. In this case, it is sufficient if the terms of a plurality of second tokens can be appropriately classified into groups.

[0032] The processor generates tokens s0, s1, ..., s t-1 , is based on the bit string based on the group classified in the process up to S106, and a group is selected based on whether the value of the subsequent bit is appropriate, 0 or 1 (S108). The processor selects a group for the second token s so that the bit string indicating the embedding information (e.g., some kind of identifier) ​​to be embedded as a digital watermark matches the bit string obtained by concatenating the numerical values ​​based on the group of each token. t Select a group.

[0033] That is, the processor generates a sequence of generated tokens s0, s1, ..., s t-1 A group indicating an appropriate value (0 or 1) to be continued (concatenated) with the bit string corresponding to the embedded information is selected. In the following, the explanation will be continued by taking an identifier as an example of embedded information.

[0034] For example, the processor selects group A if the (t+1)th bit (addendum data) from the front of the embedded information, which is a bit string (data string) indicating an identifier, is 0, and selects group B if the bit is 1.

[0035] The processor selects one of the candidates belonging to the selected group as the second token s t For example, as shown in the figure, if the next bit is set to 0, the processor selects group A, and selects candidate c2 from the candidates belonging to group A as the second token s t Select as.

[0036] In this way, the processor performs the following operations on the data string indicating the identifier: t-1 , second token s t The second token s is matched with the concatenation of the values ​​of the group to which it belongs (the embedded information is part or all of the embedded information from the front). t Select .

[0037] In other words, if the embedding information is completed by selecting one more bit, the processor selects the second token s so that the embedding information obtained by concatenating the classification results (0 or 1) of each token matches the data related to the identifier. t If it is still in the middle, select the second token s so that the embedded information matches the data related to the identifier from the front. t Select .

[0038] By repeating the process of sequentially selecting tokens starting from the 0th token so that they match the data string indicating the identifier, the data string indicating the identifier can be embedded into successive tokens. By repeatedly executing this process, the processor can embed the embedding information into the token string as a digital watermark.

[0039] As a non-limiting example, the processor may select the token in any manner after the data string (embedded information) indicating the identifier is embedded.

[0040] As a non-limiting example, after embedding the embedded information, the processor may add an error detection code or an error correction code to the embedded information by repeating the same token selection process. By adding such a code, it is possible to perform appropriate error detection or error correction even if the data is tampered with or if part of the data is destroyed by noise or the like in the transmission path. This error detection code or error correction code may have any number of digits.

[0041] The data string indicating the identifier may be predetermined to be represented by a predetermined number of digits. As another example, the processor may add a bit string indicating an end flag at the end of the embedded information or the embedded information to which a code has been added, by the above process.

[0042] As another example, the processor may repeatedly embed the embedded information or the embedded information with a code added thereto.

[0043] In this way, the processor can generate output data by concatenating each token and embedding an identifier using a value based on the group to which each token belongs.

[0044] The number of groups does not have to be two, and may be three or more. Classifying into more than two groups allows for a wider range of possibilities for expressing embedded information related to identifiers. However, classifying into more than two groups may narrow the token selection, and particularly when the entropy of the tokens that make up the output data is low, there is a possibility that the token selection room will be very limited or even eliminated. Therefore, it is desirable to set the number of groups so that balance can be maintained.

[0045] As described above, according to the information processing device of this embodiment, when data is generated using a generative model, it is possible to control the information to be embedded in the data as a digital watermark, and as a result, it is possible to embed any embedding information (bit string) into the data.

[0046] This method makes it possible to embed any type of embedded information as a digital watermark. The embedded information embedded in the data may be an identifier indicating some specific type of information. For example, it may be an identifier indicating the entity (individual or organization) that generated the data using a generative model or that instructed the generation of the data, an identifier indicating the date the data was generated using a generative model, or an identifier indicating the generative model that generated the data.

[0047] In Retrieval Augmented Generation (RAG), the identifier may be an identifier indicating the provider (individual or organization) of the information found in the information search performed when generating data using a generative model, or an identifier indicating that information.

[0048] Furthermore, the embedded information embedded in the data may be an identifier indicating information regarding the use of the data, such as information indicating that copying of the data is prohibited or information indicating the scope of use of the data. Of course, the identifier may also be a value containing any other information.

[0049] The first token for the second token does not have to be the immediately preceding token. That is, a token two tokens before the currently focused token can be selected as the first token.

[0050] Also, when the second token is determined to be token s0, that is, the first token, the processor divides the input data into tokens to generate output data, and the result is called token s M , s M-1 ,...,s -1 (M is the number of tokens in the input data), and then the token with this negative index, e.g., token s -1 This allows embedding information (bit string) to be embedded in the generated token string in order from the first token.

[0051] As another example, after generating token s0, the processor can embed information using the value related to the classification of token s1 as the starting bit based on the classification result of token s1, without classifying token s0. In this way, rather than embedding the embedded information (bit string) from the first token in the generated token string, the embedded information (bit string) can be embedded from the second token in the token string.

[0052] In the above embodiment, each token is used individually, but the present disclosure is not limited to this. The processor can perform the same processing as above on a token string consisting of a predetermined number of tokens.

[0053] 2 is a diagram showing an example in which a token string is composed of multiple tokens, such as, for example, 10 consecutive tokens, rather than a single token. As shown in this diagram, the processor groups multiple consecutive tokens into a single token string and classifies the token string into groups (0 or 1).

[0054] This allows the processor to embed, for example, one bit of information constituting a bit string indicating an identifier in each token string (each set of multiple tokens). As a result, the processor can embed identifier information in a token string formed by concatenating multiple token strings. In other words, the processor can embed one piece of embedded information in multiple token strings by embedding a bit string indicating the embedded information one bit at a time in each set of multiple tokens.

[0055] In this case, the processor generates multiple tokens together. As a non-limiting example, the processor may generate each token constituting the token sequence based on joint probability. The processor generates multiple second token sequences g t Candidates for the second token string gt can be generated based on joint probability, and the plurality of candidates can be grouped, and the second token string gt can be selected from the plurality of candidates based on information indicating the identifier.

[0056] The series of processes can be similar to those performed for one token above by performing the same processes for each token sequence. That is, in S100, the processor generates candidates for the second token sequence from a group of token sequences based on the joint probability of the tokens, generates a hash seed S from the first token sequence gt-1, calculates hash values ​​of the multiple candidates, and classifies the multiple candidates into groups, thereby achieving the same processes as those described above.

[0057] The processor can also define token sequences with negative indices in the same way as above. For example, the processor can define token s -10 From s -1 is a token sequence g -1 , ..., can be defined as:

[0058] 3 is a diagram showing another example using a token string. As shown in this diagram, multiple token strings may contain duplicated tokens. In this case, candidates can be generated and selected for each token, and a hash value can be calculated for each token string even if the grouping contains duplicated tokens. This improves the range of token selection and ultimately the robustness of the generated data.

[0059] By using multiple tokens as a token string as shown in FIG. 2 or FIG. 3, it is possible to embed appropriate information even in cases where processing each token individually results in low entropy.

[0060] In the above, the generation of tokens can be defined based on a generation model used by the information processing device.

[0061] The embedding of information using a token or a token sequence by the above-described information processing device does not require information about a generative model when reading the information. That is, the reading information processing device can properly acquire the embedded information even without acquiring information about the generative model. Furthermore, since the reading information processing device can acquire information without using a generative model, it can acquire the embedded information through high-speed calculations.

[0062] As described above, the information processing device according to this embodiment can embed embedded information in data by having the processor select tokens to be included in a token sequence from candidate tokens forming the data based on the embedded information. The processor can, for example, generate tokens forming the data using a generative model and select tokens to be included in a token sequence that is output as data. Note that if candidate tokens can be classified into groups without generating them, tokens corresponding to the embedded information may be directly selected and generated without actually generating candidate tokens. This principle can also be applied to the following description.

[0063] The processor can select each token in the token sequence based on a bit corresponding to the bit sequence that indicates embedded information.

[0064] That is, the information processing device uses a generative model to generate multiple candidates (tokens) for the second token that follows the first token that has already been generated and selected, classifies these candidates into multiple groups, and selects tokens from this classification that are appropriate for the order that matches the bit string of the embedded information.

[0065] The information processing device can also generate second token candidates from at least the first tokens that have already been generated and selected. Furthermore, the information processing device can select the second token based on at least the first token. More specifically, the information processing device can output second token candidates based on the first tokens in the generative model and / or classify the second token candidates into groups based on the first tokens.

[0066] Of course, the token into which the information is embedded may be a combination of multiple tokens (token portion) as shown in Figure 2 or 3. In other words, the information processing device may associate each bit of the bit string of the embedded information with a single token unit, or may associate each bit of the bit string of the embedded information with multiple token units. The token portion may be one token or multiple tokens.

[0067] (Embodiment relating to a reading device)

[0068] As a non-limiting example, the information processing device acquires embedded information from data in which information has been embedded using a digital watermarking technique by the above-mentioned device. As described above, the information processing device can acquire the embedded information without using information related to a generative model.

[0069] 4 is a diagram illustrating an example of embedded information acquisition according to an embodiment, in which data can be divided into tokens, and a processor can use the divided tokens to acquire the embedded information.

[0070] In addition, when one piece of embedded information is embedded in multiple token sequences, the processor can similarly acquire the embedded information by processing each token sequence shown in Figure 2, etc. In other words, the processor can acquire the information based on the data embedding conditions. The user may determine in advance the format to be used for embedding the information, or the information may be embedded in the generated data, and the processor on the reading side may extract the conditions and then execute the processing.

[0071] The processor, like S102, t-1 By using a predetermined hash seed for t The hash seed S for calculating the hash value of the message is acquired (S200). The predetermined hash seed and the algorithm for calculating the hash seed S may be the same as the hash seed and algorithm used in S102.

[0072] The processor generates a second token s using the hash seed S in the same manner as in S104. t The hash value h of t (S202). The hashing algorithm may be the same as the algorithm used in S104.

[0073] The processor, like S106, calculates the hash value h t Based on the second token s t are classified into groups (S204). The groups are the groups into which the second token candidates are classified in S106. The rules for classifying the groups may be the same as the rules used in S106.

[0074] The processor generates second tokens s based on the classified group. t The value indicated by (for example, the second token s tThen, the bit value corresponding to the group into which the image data is classified is acquired (S206).

[0075] The processor obtains the value indicated by each token by repeating the process from s0 as the second token until the final token or embedded information of the required length is obtained, and then concatenates these values ​​to obtain the embedded information (e.g., an identifier).

[0076] As described above, according to this embodiment, the information processing device can appropriately acquire information from data that the information is embedded in. The information processing device can quickly acquire information embedded in generated data without acquiring information related to the generative model used to generate the data.

[0077] (Application example)

[0078] 5 is a diagram illustrating a non-limiting example of the use of the process according to the embodiment described above. The process in this diagram may be executed, for example, in an information processing system including one or more processors installed in one or more computers and one or more storage devices (including those provided in the computers). The information processing method described above can be partially applied to this information processing system.

[0079] The information processing system is a system that generates and outputs third data, which is output data for first data, which is input data, using a first model, which is a probabilistic model trained using pre-learning data. The first model may be, for example, a base model or a generative model.

[0080] The first data is, for example, text for asking a question to the first model, and is called a prompt. Depending on the context, the first data can be interpreted as a question for which an answer is generated by the first model. In the following, the term "question" may conceptually include either a question from a user or an instruction from a user. Also, in the following, the term "question" may conceptually include either data input by a user or processed data obtained by performing predetermined processing on that data.

[0081] The first data may include any of images, audio, video, and sensor data. The third data is, for example, text indicating an answer to a question. In the following, "answer" may conceptually include either an answer to a question or a response to an instruction. Furthermore, in the following, "answer" may conceptually include either data generated by the first model or processed data obtained by performing predetermined processing on that data. Like the first data, the third data may include any of images, audio, video, and sensor data. Furthermore, both the first data and the third data are not limited to these data and may be data in other formats.

[0082] The first model is stored in a storage device within the information processing system and is formed by a processor within the information processing system referencing the storage device. The information processing system performs inference using the first model formed by the processor.

[0083] The information processing system first receives first data as input data (S10). The input of data may be performed to at least one processor in the information processing system via various interfaces. For example, the at least one processor in the information processing system may receive the first data input by a user as input data.

[0084] The first data input by the user is, for example, data input or specified by the user via the user interface of the information terminal, and may be text entered using a keyboard, text selected in a sentence, an image dragged and dropped with a mouse, or audio picked up by a microphone, etc. Here, the user's information terminal is an example of an information processing device in an information processing system, and is one element constituting the information processing system.

[0085] At least one processor in the information processing system acquires second data for performing a search necessary to generate third data based on the first data received as input data (S12). The second data is, for example, a query to be input to a search engine. Information searched using this second data is used as additional input when determining the output of the third data. In other words, the second data is data for searching for information related to an answer to the first data.

[0086] At least one processor in the information processing system may, for example, tokenize the first data by inputting the first data into a trained first model, and acquire a query required for a search as the second data based on the acquired tokens. As a non-limiting example, the first model may be a large-scale language model, and in this case, the processor acquires the second data by extracting a search query for generating third data from the tokenized first data for the input first data indicating a question or the like.

[0087] Furthermore, at least one processor in the information processing system may, for example, vectorize the first data and acquire the acquired vector as second data, which is a query required for a search. Furthermore, at least one processor in the information processing system may, for example, acquire the first data itself as second data, which is a query required for a search. Furthermore, at least one processor in the information processing system may use a conventional method to acquire second data, which is a query for a search engine described below, based on the first data.

[0088] At least one processor in the information processing system uses the acquired second data as a query to search for information required to generate the third data using a search engine (S14). The search engine is stored in a storage device in the information processing system, similar to the first model, and is formed by at least one processor in the information processing system referencing this storage device. The information processing system performs a search for the required information using the search engine formed by the processor.

[0089] The information processing system inputs the second data into a search engine and acquires information to be used to generate the third data from the data group (S16). The processing of the search engine may be executed by at least one processor in the information processing system. As a non-limiting example, this processing enables the information processing system to acquire information to be used to generate an answer to the first data from the data group based on the first data.

[0090] The processing by the search engine can be performed by, for example, using the data to be searched, which is information within a data group, as a key to extract a key for the query. Note that the search for the key by the search engine is not limited to a specific method, and any general method for obtaining a key for a query can be used.

[0091] The key and the query may be encoded. For example, the search engine may execute a search using a function defined by the encoded key and the encoded query. The key and query encoding may also be pre-trained during pre-training of the first model.

[0092] As a non-limiting example, if the first model is a large-scale language model, at least one processor in the information processing system can generate a query required for the search from the tokenized text and use this query to reference keys in the data group to extract information that may include text required to answer the input text.

[0093] Furthermore, the search engine may be an engine formed by any method, such as a model formed by keyword matching, an attention mechanism, or a model formed by a learning method using reinforcement learning, etc. However, it is not limited to these, and an engine formed by other methods may also be used.

[0094] At least one processor in the information processing system reflects the search results of the search engine in generating output data in the first model (S18).

[0095] At least one processor in the information processing system generates third data based on the input data received in S10 and the information obtained by the search to generate third data, which is output data in the first model (S20). That is, the at least one processor in the information processing system can use this information to generate an answer to the first data using the first model.

[0096] In the process of S20, the information processing device according to the present disclosure described in the above embodiment generates the third data while embedding embedded information. The embedded information is, for example, an identifier as information about the provider of the information searched for in S16.

[0097] For example, when generating the third data, the information processing device acquires multiple candidates for the second token from the first model. The information processing device executes an embedding process based on the multiple candidates and the process of the above-described embodiment, thereby enabling the information processing device to embed, for example, an identifier indicating the information provider in the third data.

[0098] The information processing device that generates the third data can use the tokens of the first data when generating the first one or more tokens of the third data. In this case, the information processing device that acquires embedded information from the third data can acquire embedded information such as bit values ​​assigned to the first one or more processor tokens of the third data by using the first data.

[0099] Of course, in this application example as well, the information processing device can execute processing not for each token, but for each token string formed from a plurality of tokens.

[0100] At least one processor in the information processing system may output the generated third data (answer) to the information terminal of the user who input the first data in S10. The output third data (answer) with the embedded information embedded therein may be, for example, displayed on the user's information terminal, read aloud on the information terminal, printed via the information terminal, or the like.

[0101] For example, at least one processor in the information processing system can input input data and information obtained by a search into a first model to generate a response or answer using this information in addition to knowledge based on pre-trained data as third data with embedded information embedded therein. As an example, at least one processor in the information processing system can process the input data and information obtained by a search into a prompt in a predetermined format and input this prompt into the first model to generate the third data.

[0102] As another non-limiting example, at least one processor in the information processing system may input information obtained by the search into the first model as input to the first model's internal calculations, and input the input data as a prompt to the first model to generate third data. In this example, the information input point is appropriately designed depending on the architecture of the first model, and may be, for example, any of the input layer, intermediate layer, and output layer of the first model.

[0103] The first model can generate the third data using information about the encoded key set in the data group. As a non-limiting example, when the first model is a large-scale language model, the first model obtains information as a search result for input text as an encoded embedding vector, and reflects the information of this encoded embedding vector in output data, thereby generating more accurate text, etc. as an answer to a question, etc.

[0104] According to this information processing system, an identifier can be embedded in the output data as a response. This embedded data is, for example, an identifier that indicates the provider of the information, as described above.

[0105] At least one processor in the information processing system may calculate and assign a contribution level for information obtained from the data group and used to generate the third data (S22). Note that in the present disclosure, the process of S22 is not essential, but is shown as an example of application.

[0106] The contribution level is an example of an index for evaluating the usefulness of information. The contribution level may be assigned at the time when the information is provided to the data group. By this process, the information processing system can, as a non-limiting example, calculate the contribution level of the information in generating an answer using the first model.

[0107] According to the embedded information disclosed herein, it is possible to obtain an identifier of an information provider from output data, and to assign a contribution level to the information provider based on this identifier. This embedded information will not be replaced with another identifier, even if, for example, part of the token is tampered with or deleted. Furthermore, because the embedded information is hashed, a process that is generally difficult to analyze, it is difficult for a tamperer to rewrite it with an arbitrary identifier through tampering or the like.

[0108] As described above, this application example is shown as one example and is not limiting, and embedding of information according to the present disclosure does not exclude suitable applications other than this application example.

[0109] Some or all of the devices (information processing devices) in the above-described embodiments may be configured as hardware, or may be configured as software (programs) executing information processing by a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), etc. When software information processing is configured, software that realizes at least some of the functions of each device in the above-described embodiments may be stored on a non-transitory storage medium (non-transitory computer-readable medium) such as a CD-ROM (Compact Disc-Read Only Memory) or a USB (Universal Serial Bus) memory, and the software information processing may be executed by loading the software into a computer. The software may also be downloaded via a communications network. Furthermore, all or part of the software processing may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), thereby allowing the software information processing to be executed by hardware.

[0110] The storage medium that stores the software may be a removable medium such as an optical disk, or a fixed medium such as a hard disk or memory. The storage medium may be located inside the computer (such as a main memory or auxiliary memory), or may be located outside the computer.

[0111] 6 is a block diagram showing an example of the hardware configuration of each device (information processing device) in the above-described embodiment. Each device may be realized as a computer 7 including, for example, a processor 71, a main storage device 72 (memory), an auxiliary storage device 73 (memory), a network interface 74, and a device interface 75, all of which are connected via a bus 76.

[0112] Although the computer 7 in FIG. 6 includes one of each component, it may include multiple of the same component. Also, while FIG. 6 shows one computer 7, the software may be installed on multiple computers, and each of the multiple computers may execute the same or different parts of the software. In this case, a distributed computing configuration may be used in which each computer communicates with the other computers via a network interface 74 or the like to execute the processing. In other words, each device (information processing device) in the above-described embodiment may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize its functions. Furthermore, the system may be configured such that information sent from a terminal is processed by one or more computers located on a cloud, and the processing results are sent to the terminal.

[0113] The various calculations of each device (information processing device) in the above-described embodiments may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, the various calculations may be distributed to multiple processor cores within a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. disclosed herein may be implemented by at least one processor and storage device provided on a cloud that can communicate with computer 7 via a network. Thus, each device in the above-described embodiments may be implemented in the form of parallel computing using one or more computers.

[0114] The processor 71 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, etc.) that at least controls or performs calculations on a computer. The processor 71 may also be a general-purpose processor, a dedicated processing circuit designed to perform a specific calculation, or a semiconductor device that includes both a general-purpose processor and a dedicated processing circuit. The processor 71 may also include an optical circuit or a calculation function based on quantum computing.

[0115] The processor 71 may perform arithmetic processing based on data or software input from each device, etc., configured inside the computer 7, and may output arithmetic results or control signals to each device, etc. The processor 71 may control each component constituting the computer 7 by executing the OS (Operating System) of the computer 7, applications, etc.

[0116] Each device (information processing device) in the above-described embodiments may be realized by one or more processors 71. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other via wire or wirelessly.

[0117] The main memory device 72 may store instructions executed by the processor 71 and various data, etc., and information stored in the main memory device 72 may be read by the processor 71. The auxiliary memory device 73 is a memory device other than the main memory device 72. Note that these memory devices refer to any electronic component capable of storing electronic information and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory device for saving various data, etc. in each device (information processing device) in the above-described embodiments may be realized by the main memory device 72 or the auxiliary memory device 73, or may be realized by an internal memory built into the processor 71. For example, the memory in the above-described embodiments may be realized by the main memory device 72 or the auxiliary memory device 73.

[0118] When each device (information processing device) in the above-described embodiments is configured with at least one storage device (memory) and at least one processor connected (coupled) to this at least one storage device, at least one processor may be connected to one storage device. At least one storage device may be connected to one processor. A configuration in which at least one processor among multiple processors is connected to at least one storage device among multiple storage devices may also be included. This configuration may also be realized by storage devices and processors included in multiple computers. Furthermore, a configuration in which a storage device is integrated with a processor (for example, a cache memory including an L1 cache and an L2 cache) may also be included.

[0119] The network interface 74 is an interface for connecting to the communication network 8 wirelessly or via a wire. The network interface 74 may be an appropriate interface, such as one that conforms to an existing communication standard. The network interface 74 may exchange information with an external device 9A connected via the communication network 8. The communication network 8 may be any one of a WAN (Wide Area Network), a LAN (Local Area Network), a PAN (Personal Area Network), etc., or a combination thereof, as long as information is exchanged between the computer 7 and the external device 9A. An example of a WAN is the Internet, an example of a LAN is IEEE 802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication), etc.

[0120] The device interface 75 is an interface such as USB that directly connects to the external device 9B.

[0121] The external device 9A is a device connected to the computer 7 via a network. The external device 9B is a device directly connected to the computer 7.

[0122] For example, the external device 9A or the external device 9B may be an input device. The input device may be a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the computer 7. Alternatively, the external device 9A or the external device 9B may be a device equipped with an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0123] Furthermore, the external device 9A or the external device 9B may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel, or a speaker that outputs sound or the like. Alternatively, the external device 9A or the external device 9B may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0124] Furthermore, the external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage such as an HDD.

[0125] Furthermore, the external device 9A or the external device 9B may be a device having some of the functions of the components of each device (information processing device) in the above-described embodiment. That is, the computer 7 may transmit some or all of the processing results to the external device 9A or the external device 9B, or may receive some or all of the processing results from the external device 9A or the external device 9B.

[0126] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, a-b, a-c, b-c, or a-b-c. It may also include multiple instances of any element, such as a-a, a-b-b, a-a-b-b-c-c, etc. It also includes the addition of elements other than the listed elements (a, b, and c), such as having d, as in a-b-c-d.

[0127] In this specification (including claims), when expressions such as "using / using data as input / based on / according to / in response to" (including similar expressions) are used, unless otherwise specified, this includes cases where the data itself is used, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of data, etc.) is used. Furthermore, when a statement is made that a result is obtained "using data as input / based on / according to / in response to" (including similar expressions), this includes cases where the result is obtained based solely on the data, or where the result is influenced by other data, factors, conditions, and / or states other than the data, unless otherwise specified. Furthermore, when a statement is made that "data is output" (including similar expressions), this includes cases where the data itself is used as output, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of data, etc.) is used as output, unless otherwise specified.

[0128] When the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms that encompass any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, and physically connection / coupling. These terms should be interpreted appropriately according to the context in which they are used, but any connection / coupling form that is not intentionally or naturally excluded should be interpreted as being included in these terms without limitation.

[0129] In this specification (including the claims), the expression "A configured to B" may include the physical structure of element A having a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A being configured / set to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, if element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached to it.

[0130] When used in this specification (including the claims), terms implying containing or possessing (e.g., "comprising" or "including" and "having"), they are intended to be open-ended terms that include containing or possessing things other than the object designated by the object of the term. When the object of such terms implies no quantity or a singular number (e.g., expressions using the articles "a" or "an"), the expression should be construed as not being limited to a specific number.

[0131] In this specification (including the claims), although expressions such as "one or more" or "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles a or an) are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles a or an) should be interpreted as not necessarily being limited to a specific number.

[0132] In this specification, when a particular advantage / result is described as being obtained from a particular configuration of an embodiment, it should be understood that the same advantage / result can also be obtained from one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or circumstances, and that the effect is not necessarily obtained by the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or circumstances are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same or a similar configuration.

[0133] When terms such as "maximize" and "maximization" are used in this specification (including the claims), they include finding a global maximum, finding an approximation of a global maximum, finding a local maximum, and finding an approximation of a local maximum, and should be interpreted accordingly according to the context in which the term is used. They also include finding approximations of these maxima probabilistically or heuristically. Similarly, when terms such as "minimize" and "minimization" are used, they include finding a global minimum, finding an approximation of a global minimum, finding a local minimum, and finding an approximation of a local minimum, and should be interpreted accordingly according to the context in which the term is used. They also include finding approximations of these minima probabilistically or heuristically. Similarly, when terms such as "optimize" and "optimization" are used, they include finding a global optimum, finding an approximation of a global optimum, finding a local optimum, and finding an approximation of a local optimum, and should be interpreted accordingly according to the context in which the term is used. It also includes finding approximations of these optimum values ​​probabilistically or heuristically.

[0134] In this specification (including claims), when multiple pieces of hardware perform a predetermined process, the pieces of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Furthermore, some of the hardware may perform part of the predetermined process, and other hardware may perform the rest of the predetermined process. In this specification (including claims), when expressions such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" (including similar expressions) are used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. Note that hardware may also include an electronic circuit or a device including an electronic circuit.

[0135] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices may store only a portion of the data, or may store the entire data. Also, a configuration in which only some of the multiple storage devices store data may be included.

[0136] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, and partial deletions are possible within the scope of the conceptual idea and spirit of the present disclosure, which is derived from the content defined in the claims and their equivalents. For example, when numerical values ​​or formulas are used in the above-described embodiments, they are shown for illustrative purposes and do not limit the scope of the present disclosure. Furthermore, the order of each operation shown in the embodiments is also illustrative and does not limit the scope of the present disclosure.

Claims

1. An information processing device comprising: one or more memories; and one or more processors, wherein the one or more processors select tokens to be included in a token string to be output as data from among tokens generated by a generative model based on embedding information, and embed the embedded information in the data.

2. An information processing device comprising one or more memories and one or more processors, wherein the one or more processors cause a generative model to output a token sequence in which the same embedding information is repeatedly embedded.

3. An information processing device according to claim 1 or 2, wherein each of the token parts of the token sequence is selected from tokens generated by the generative model based on a corresponding bit of a bit sequence indicating the embedded information.

4. The information processing device according to claim 3, wherein the one or more processors classify tokens generated by the generative model as candidates for tokens in the second token portion of the token string into a plurality of groups, and select a token to be output as the second token portion from the tokens in the group corresponding to a bit to be associated with the second token portion.

5. The information processing device according to claim 4, wherein the one or more processors classify tokens generated by the generative model as candidates for tokens in a second token portion of the token string into a plurality of groups based on a first token portion preceding the second token portion.

6. The information processing device according to claim 4, wherein the one or more processors generate candidate tokens for the second token part using the generative model based on a first token part preceding the second token part, and classify the tokens generated by the generative model as candidate tokens for the second token part into the multiple groups based on the token selected as the first token part.

7. An information processing device comprising: one or more memories; and one or more processors, wherein for a token sequence formed by concatenating a predetermined number of tokens, the one or more processors convert each of a plurality of second token sequence candidates generated based at least on a first token sequence generated in the past according to a predetermined rule, classify the plurality of candidates into groups based on the conversion results, select the second token sequence from the plurality of candidates such that data obtained by concatenating embedded information and candidates for additional data indicated by each group of the plurality of candidates becomes a data sequence related to embedded information, concatenate the second token sequence to output data, and concatenate the additional data related to the second token sequence to the embedded information.

8. The information processing device according to claim 7, wherein the specified rule is a transformation including calculation of a hash value, and the one or more processors transform the multiple candidates based on a specified hash seed and the first token sequence.

9. The information processing device according to claim 8, wherein the one or more processors convert the plurality of candidates based on a hash value of the first token string converted using the predetermined hash seed.

10. The information processing device according to claim 9, wherein the one or more processors obtain a hash seed based on a hash value of the first token sequence converted using the specified hash seed, and convert the multiple candidates using the hash seed.

11. The information processing device according to claim 7, wherein the one or more processors divide the plurality of candidates into a group indicating 0 and a group indicating 1.

12. The information processing device according to claim 11, wherein the one or more processors select the second token sequence so that a data sequence obtained by concatenating the embedded information with the candidates for the additional data indicated by each of the multiple candidate groups becomes part or all of the data sequence indicating the embedded information.

13. The information processing device according to claim 12, wherein the part of the data string representing the embedded information is data that matches a beginning of the data string representing the embedded information.

14. The information processing device according to claim 12, wherein the one or more processors select the second token sequence so that the concatenated data sequence matches a bit sequence indicating the identifier, either currently or in the future.

15. An information processing device according to any one of claims 7 to 14, wherein when the token sequence is a token sequence formed from a plurality of tokens, the one or more processors generate the second token sequence based on the joint probability of the plurality of tokens that form the second token sequence.

16. The information processing device according to claim 7, wherein the one or more processors generate candidates for the second token sequence using a generative model.

17. The information processing device according to claim 7, wherein the one or more processors: after the embedded information is generated as a data sequence related to the identifier, generate an error detection code or an error correction code for the data sequence related to the identifier; convert multiple second token sequence candidates generated based at least on a first token sequence generated in the past according to a predetermined rule; classify the multiple candidates into groups based on the conversion results; select the second token sequence into which the additional data that becomes the error detection code or the error correction code is converted from the second token sequence candidates; concatenate the second token sequence to output data; and concatenate the additional data related to the second token sequence to the embedded information.

18. The information processing device according to claim 7, wherein the embedded information is an identifier indicating a predetermined type of information.

19. An information processing device comprising: one or more memories; and one or more processors, wherein the one or more processors divide data into a token string; identify bits corresponding to each token portion contained in the token string; and obtain a bit string obtained by concatenating each bit corresponding to each token portion as embedded information embedded in the data.

20. The information processing apparatus of claim 18, wherein the one or more processors identify bits corresponding to a next token portion based on a previous token portion.

21. An information processing device comprising: one or more memories; and one or more processors, wherein the one or more processors divide data into a token sequence formed by concatenating a predetermined number of tokens, which is one or more numbers; converting the token sequence according to a predetermined rule based on a token sequence preceding the token sequence; classifying the token sequence into groups based on the conversion result; obtaining a value based on the group; concatenating the values ​​of each of the token sequences; and obtaining embedded information embedded in the data from the concatenated result.

Citation Information

Patent Citations

  • Tamper-resistant text stream watermarking

    US20070047758A1