Generative content risk prediction method and device, storage medium and electronic equipment

By extracting the internal representation of the intermediate network level of the large language model and using the target classifier to detect risks, the risk challenges in the large language model when generating content is solved, and efficient risk prediction and content security management are achieved.

CN120046153APending Publication Date: 2025-05-27ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411941079.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The risks and challenges faced by large language models when generating content, including potential risks in harmlessness and honesty, lead to sensitive information leakage and security issues.

Method used

By extracting the internal representation of the large language model at the intermediate network level during the content generation process, and using the target classifier to determine the probability of generating risk content, detect and predict the risk of generating content in real time.

Benefits of technology

It realizes the identification and processing of risky content in the early stages of content generation, improves the overall security of generated content, avoids users being affected by bad content, and significantly improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046153A_ABST
    Figure CN120046153A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a generative content risk prediction method and device, a storage medium and electronic equipment.The method comprises the steps that firstly, a risk problem text is obtained, then the risk problem text is input into a large language model for content generation, and internal representation of a middle network hierarchy of the large language model in the content generation process is extracted; and finally, based on the internal representation, determining the risk content output probability of the large language model in the content generation process by utilizing a target classifier. By predicting generation of the risk content in advance and performing risk management and control in time, the risk can be prevented from being exposed to the user, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a generative content risk prediction method, device, storage medium and electronic device. Background Art

[0002] As the scale of machine learning models continues to expand, large models are increasingly used in finance, medicine, autonomous driving and other fields, but they also bring great risks and challenges. Taking large language models as an example, they are mainly reflected in two aspects: harmlessness and honesty. Among them, harmlessness means that the content generated by the large language model does not contain information such as bias, discrimination, and personal privacy, and honesty means that the content generated by the large language model does not contain false, forged, and fraudulent content.

[0003] Given the risk challenges faced by large language models, it is particularly important to perform risk detection on large language models. By performing risk detection on large language models, we can evaluate and identify potential risks when large language models generate content, thereby triggering corresponding security defense modules to handle and prevent risks from occurring. Otherwise, large language models may expose sensitive information, leak private data, etc., thus causing a series of security issues. At present, there is an urgent need to provide a new risk detection solution to further improve the risk prevention and control performance of large language models. Summary of the invention

[0004] The embodiment of this specification provides a generative content risk prediction method, which can avoid exposing risks to users and improve user experience by predicting the generation of risky content in advance and performing risk management in a timely manner. The generative content risk prediction method includes:

[0005] Get the risk question text;

[0006] Inputting the risk question text into a large language model for content generation, and extracting the internal representation of the intermediate network layer of the large language model in the content generation process;

[0007] Based on the internal representation, a target classifier is used to determine the probability that the large language model outputs risky content during content generation.

[0008] Further, in some embodiments, the intermediate network level includes at least one hidden layer in the middle position of the large language model network structure, and the generated content of the large language model includes a word sequence;

[0009] The extracting the internal representation of the intermediate network layer of the large language model in the content generation process includes:

[0010] Extracting output vectors of each hidden layer during the content generation process;

[0011] The output vectors are concatenated to obtain the internal representation.

[0012] Furthermore, in some implementations, extracting the output vector of each hidden layer during the content generation process includes:

[0013] The output vector of each hidden layer when processing the risk question text and not generating word-units is extracted.

[0014] Furthermore, in some implementations, extracting the output vector of each hidden layer during the content generation process includes:

[0015] Extracting the output vector corresponding to the target word in the currently generated word sequence from each of the hidden layers;

[0016] The target word is the last word in the currently generated word sequence.

[0017] Further, in some implementations, determining the probability of the large language model outputting risky content during content generation using a target classifier based on the internal representation includes:

[0018] Encoding the internal representation using the target classifier to obtain target features;

[0019] The probability that the large language model outputs risky content during content generation is determined according to the target feature.

[0020] Further, in some embodiments, the target classifier includes a linear transformation layer;

[0021] The step of encoding the internal representation using the target classifier to obtain target features includes:

[0022] The linear transformation layer is used to perform weighted summation on the internal representation, and the target feature is calculated based on the summation result and the bias term.

[0023] Further, in some embodiments, the target classifier further includes a nonlinear activation function;

[0024] The determining, according to the target feature, the probability that the large language model outputs risky content during content generation includes:

[0025] The target feature is converted into the probability of the large language model outputting risky content during the content generation process using the nonlinear activation function.

[0026] Furthermore, in some embodiments, the method further comprises:

[0027] When the probability that the large language model outputs risky content during the content generation process is greater than or equal to a preset risk probability, the large language model is controlled to stop generating content.

[0028] Furthermore, in some embodiments, the method further comprises:

[0029] When the probability that the large language model outputs risky content during the content generation process is less than a preset risk probability, generating a new word by using the large language model;

[0030] The new word-gram is used to update the current word-gram sequence and the internal representation of the intermediate network layer of the large language model during the content generation process, so as to determine the probability of the large language model outputting risky content when generating the next word-gram based on the updated internal representation and using the target classifier.

[0031] Furthermore, in some embodiments, the method further comprises:

[0032] Obtain a training dataset containing risk labels;

[0033] Generate an internal representation of each sample data in the training data set using the trained large language model;

[0034] Inputting the internal representation into a preset neural network to obtain the risk prediction probability of each sample data;

[0035] The loss function is calculated based on the risk prediction probability of each sample data and the corresponding risk label, and the model parameters in the preset neural network are iteratively adjusted according to the calculated loss value. When the iteration termination condition is met, the target classifier is obtained.

[0036] The embodiment of this specification also provides a generative content risk prediction device, the device comprising:

[0037] Risk problem acquisition module, used to obtain risk problem text;

[0038] An internal representation determination module, used for inputting the risk question text into a large language model for content generation, and extracting the internal representation of the intermediate network layer of the large language model during the content generation process;

[0039] A risky content prediction module is used to determine the probability of the large language model outputting risky content during the content generation process based on the internal representation and using a target classifier.

[0040] The embodiments of the present specification also provide a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.

[0041] An embodiment of the present specification also provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.

[0042] The embodiments of the present specification also provide a computer program product, wherein the computer program product stores at least one instruction, and the at least one instruction is suitable for being loaded by a processor and executing the above method steps.

[0043] In the embodiment of the present specification, the risk problem text is first obtained, and then the risk problem text is input into the large language model for content generation, and the internal representation of the intermediate network layer of the large language model in the content generation process is extracted. Finally, based on the internal representation, the target classifier is used to determine the probability of the large language model outputting risk content in the content generation process. On the one hand, by detecting the risk probability in real time during the process of generating content by the large language model, it is possible to instantly understand whether the generated content is risky, and to identify and process risky content in the early stage of content generation, with higher efficiency and accuracy; on the other hand, by detecting the risk probability of generated content in real time, the overall security of the generated content can be effectively improved, which helps to prevent risky content from appearing in the final generated results and avoid users from being affected by bad content; moreover, by reducing the generation of bad content, the user experience can be significantly improved; on the other hand, based on the feedback of the risk probability, the generation strategy of the large language model can be dynamically adjusted. For example, when it is detected that the risk probability of the generated content is high, the content generation process can be interrupted or the model parameters can be adjusted to reduce the risk probability, making the generation process more intelligent and controllable. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of a system architecture for applying the embodiments of this specification.

[0045] Figure 2 A flowchart of a generative content risk prediction method provided in an embodiment of this specification.

[0046] Figure 3 A flowchart of another generative content risk prediction method provided in an embodiment of this specification.

[0047] Figure 4 A schematic diagram of the principle of an application-generated content risk prediction method provided in an embodiment of this specification.

[0048] Figure 5 A schematic diagram of the structure of a generative content risk prediction device provided in an embodiment of this specification.

[0049] Figure 6A schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0051] Figure 1 A schematic diagram of an architecture to which the embodiments of this specification can be applied is shown.

[0052] like Figure 1 As shown, the system architecture 100 may include one or more of terminal devices such as a smart phone 101, a portable computer 102, a desktop computer 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. The terminal device may be various electronic devices with text data processing functions, and the electronic device may have a display screen, which is used to display risk question texts, text content output by a large language model, etc.

[0053] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.

[0054] The generative content risk prediction method provided in the embodiments of this specification is generally executed by a terminal device, and accordingly, the generative content risk prediction device is generally set in the terminal device. However, it is easy for those skilled in the art to understand that the generative content risk prediction method provided in the embodiments of this specification can also be executed by the server 105, and accordingly, the generative content risk prediction device can also be set in the server 105, and this exemplary embodiment does not make special limitations on this.

[0055] See also Figure 2, which is a flow chart of a generative content risk prediction method according to an embodiment of this specification. The generative content is content automatically generated by a machine learning model, and the content includes but is not limited to text, images, audio or video. In an embodiment of this specification, the generative content risk prediction method is applied to a generative content risk prediction device or an electronic device equipped with a generative content risk prediction device. Figure 2 The process shown in the figure is described in detail, and the generative content risk prediction method may specifically include the following steps:

[0056] S202, obtaining risk problem text;

[0057] In one or more embodiments of the present specification, risk problem text refers to prompt text containing potential risks, which is subsequently used to input into a large language model for content generation, and combined with a target classifier to determine whether the large language model will output risky content, thereby testing the risk management performance of the large language model.

[0058] Among them, the potential risks in the risk question text include but are not limited to the risk content such as personal privacy, false information, and violent information output by the large language model. The embodiment of this specification does not limit the definition of the risk question text, and different risk contents can be customized according to specific scenarios and content generation requirements.

[0059] The risk problem text in the embodiments of this specification can directly use text data. Of course, it can also be converted from non-text data such as voice, image, table, etc. into text data, so as to obtain the risk problem text for risk detection. For example, voice is converted into text using voice recognition technology. For another example, image recognition technology is used to generate a description text of an image. Then, the converted text data can be input into a large language model for risk prediction and assessment.

[0060] S204, inputting the risk problem text into a large language model for content generation, and extracting the internal representation of the intermediate network layer of the large language model in the content generation process;

[0061] In one or more embodiments of this specification, the large language model may adopt a Transformer architecture, or an architecture that is a variant and extension of the Transformer architecture, such as BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), T5 (Text-to-Text Transfer Transformer), and other architectures. Other architectures such as recurrent neural networks and convolutional neural networks may also be used. The embodiments of this specification do not specifically limit the architecture type of the large language model. It should be noted that the large language model has been trained on a large amount of data and can perform tasks such as text generation or text understanding.

[0062] Before inputting the risk problem text into the large language model, the risk problem text can be preprocessed according to the actual application requirements, such as removing irrelevant information, normalizing punctuation, etc. Then, the preprocessed risk problem text will be converted into a form suitable for model input, such as encoding into a word sequence through a vocabulary or directly converting it into a word embedding.

[0063] Take the Transformer-based large language model as an example. This model can process sequence data and capture long-distance dependencies. When the large language model starts to generate content, it can predict the next most likely word based on the previous words and the language patterns learned by the model.

[0064] It should be noted that in the content generation process of the embodiments of this specification, the internal representation of the intermediate network level of the large language model can be extracted. The intermediate network level includes at least one hidden layer in the middle position of the large language model network structure. Accordingly, the internal representation refers to the state of each hidden layer, which is used to reflect the understanding of the large language model of the input risk question text and how to construct the context to generate subsequent words.

[0065] For example, if the large language model network structure includes 5 hidden layers, the selected intermediate network layer may be the third hidden layer. In this case, the extracted internal representation is the output vector of the third hidden layer.

[0066] Optionally, the output of the intermediate network layers of a large language model is usually a high-dimensional vector. In order to make the high-dimensional data easier to understand, the feature distribution in the intermediate network layers can be displayed after dimensionality reduction through visualization techniques such as t-distributed neighborhood embedding or principal component analysis, thereby obtaining the internal representation of the intermediate network layers in the content generation process.

[0067] The embodiments of this specification do not limit the specific implementation of extracting the internal representation of the intermediate network layer of the large language model during the content generation process, as long as the internal representation of the intermediate network layer during the content generation process can be used to predict the risk content of the large language model. In the embodiments of this specification, it is important to enhance the ability to predict risks by using the information of the intermediate layer of the large language model.

[0068] S206: Based on the internal representation, determine the probability of the large language model outputting risky content during the content generation process using a target classifier.

[0069] It is understandable that when the large language model violates harmlessness and honesty when generating content, that is, when the large language model outputs risky content that violates harmlessness and honesty, it indicates that the large language model may have risk problems. For example, risky content can be any form of bad content such as sensitive information, inappropriate speech, and biased expression.

[0070] In one or more embodiments of the present specification, a large language model can be probed during the content generation process to learn the internal representation of the large language model. Probing detection refers to using a classifier to evaluate whether the representation of the internal layer of the large language model contains specific linguistic features or other information. When the large language model generates text content, Probing detection intervenes immediately before or after each token is generated to promptly evaluate whether the token is likely to contain potential risk content. Tokens are the basic units in the content generation process and can be generated one by one by the large language model according to the context.

[0071] Optionally, the target classifier refers to a binary classifier inserted in the large language model, which is used to explore the information captured at different levels within the large language model. The target classifier can be a simple logistic regression model, a support vector machine, a decision tree, a random forest, or other more complex deep learning models (such as neural networks), etc., which is not limited in this embodiment of the specification.

[0072] Specifically, the target classifier can be trained by using auxiliary tasks, that is, the target classifier can be trained to complete specific auxiliary tasks based on the internal representation of the large language model. The auxiliary tasks can be part-of-speech tagging, risk prediction, etc., and specific training tasks can be set according to actual needs.

[0073] When the auxiliary task is risk prediction, the trained target classifier can be integrated into the content generation process of the large language model. For example, the internal representation of the intermediate network layer of the large language model in the content generation process is used as the input of the trained target classifier, and a scalar or vector representing the probability of outputting risky content is output.

[0074] If the target classifier detects a risk, you can choose to discard the currently generated results and try to regenerate them, or have a manual reviewer intervene to decide whether to publish the content. In addition, the target classifier can be updated regularly to adapt to new risk types. For example, collect training sample data that represents the latest risk trends, and use the new training sample data to retrain the model to ensure that the model can accurately identify the latest risk content. By regularly checking the performance of the target classifier and adjusting the training strategy according to actual needs, the target classifier can be effectively used to monitor and reduce the probability of large language models generating risky content.

[0075] The embodiments of this specification can instantly understand whether there are risks in the generated content by detecting the risk probability in real time during the process of generating content with a large language model, and can identify and process risky content in the early stage of content generation, with higher efficiency and accuracy. By detecting the risk probability of generated content in real time, the overall security of the generated content can be effectively improved, which helps prevent risky content from appearing in the final generated results and avoids users from being affected by bad content; moreover, by reducing the generation of bad content, the user experience can be significantly improved; in addition, based on the feedback of risk probability, the generation strategy of the large language model can be dynamically adjusted. For example, when it is detected that the risk probability of the generated content is high, the content generation process can be interrupted or the model parameters can be adjusted to reduce the risk probability, making the generation process more intelligent and controllable.

[0076] See also Figure 3 , provides a flow chart of another generative content risk prediction method for the embodiment of this specification. Among them, the intermediate network level includes at least one hidden layer in the middle position of the large language model network structure. The generative content risk prediction method can specifically include the following steps:

[0077] S302, obtaining risk problem text;

[0078] For step S302, please refer to the detailed description of step S202 in another embodiment of this specification, which will not be repeated here.

[0079] The risk question text will be provided as input to the large language model in order to observe the behavior of the large language model when processing the risk question text.

[0080] S304, inputting the risk problem text into a large language model for content generation, extracting output vectors of each hidden layer in the content generation process, and concatenating the output vectors to obtain an internal representation;

[0081] In one or more embodiments of the present specification, the risk problem text is input into the large language model for content generation. When the large language model has not yet generated a word unit or has generated the first few words and has not yet generated risk content, the risk probability generated by the large language model can be predicted in advance to ensure the security of the content output by the large language model.

[0082] Optionally, in the case where the large language model has not yet generated a word-gram, the output vector of each hidden layer when processing the risk problem text and not generating a word-gram can be extracted, and the output vectors of each hidden layer can be concatenated to obtain an internal representation, so as to predict the risk probability of the large language model when generating the first word-gram based on the internal representation.

[0083] Optionally, for the case where the large language model generates the first few words and has not yet generated risk content, the output vector corresponding to the target word in the currently generated word sequence can be extracted from each hidden layer, and the output vectors of each hidden layer can be spliced ​​to obtain an internal representation, so as to predict the risk probability of the large language model in generating the next word based on the internal representation. The target word in the embodiment of this specification can be the word at the end of the currently generated word sequence. The output vector corresponding to the target word contains all the information of the current word sequence, because the target word is generated based on all the previous words.

[0084] Among them, the internal representation can help understand how the large language model gradually processes the input text, and which hidden layer representations are most critical to the final output.

[0085] In one embodiment of the present specification, risk prediction can be performed using the internal representation of the intermediate network layer of the large language model network structure through the Probing detection method, that is, l∈[1,n-1], where l is the index of the intermediate layer of the large language model network structure, and n is the total number of intermediate layers.

[0086] See also Figure 4 , the risk problem text 401 is input into the large language model 402. During the content generation process, the internal representation of at least one hidden layer 4021 in the middle position of the large language model 402 when processing the risk problem text 401 is extracted, that is, the output vector 403 of at least one hidden layer 4021 is extracted.

[0087] For example, the output vector of the risk problem text 401 or the target word in the i-th hidden layer can be extracted and recorded as emb i In order to enhance the effect of risk prediction, the output vectors of multiple hidden layers in the middle of the network structure can be selected for concatenation, namely:

[0088] emb = concat(emb i , …, emb j) (1)

[0089] Among them, i and j represent the index of the output vector of each hidden layer, that is, emb i represents the output vector of the i-th hidden layer, emb j represents the jth hidden layer output vector, concat() represents the concatenation of the output vectors of multiple hidden layers, and emb represents the concatenated new vector, which contains all the feature information of the selected hidden layers. emb corresponds to Figure 4 The output vector 403 in .

[0090] When splicing the output vectors of multiple hidden layers, the output vectors of each hidden layer can be directly connected in series to obtain a new vector, retaining the characteristic information of each layer; the output vectors of each hidden layer can also be weighted averaged, and the importance of different layers can be determined by using a gating mechanism, an attention mechanism, etc., and combined, which is not limited in the embodiments of this specification. By splicing the output vectors of multiple levels, a more comprehensive risk content feature can be obtained, which helps to improve the accuracy of risk prediction in the future.

[0091] The Probing detection method can then be used to make risk predictions using the internal representation of the intermediate network layers of the large language model network structure. Figure 4 The output vector 403 in is subjected to Probing detection, thereby obtaining the probability 405 that the large language model 402 outputs risky content. If it is determined that there is no risk based on the probability 405 of outputting risky content, a new word 404 can be generated.

[0092] When processing the input risk question text, the large language model will go through multiple hidden layers for feature extraction and transformation, where each hidden layer will output the corresponding internal representation. Importantly, the output vector of the hidden layer in the middle contains rich semantic information and context-related features, which is more original and diverse than the information of the final output layer. Therefore, it can provide more detailed insights, thereby identifying possible risks in the early stages of the content generation process.

[0093] S306, using a target classifier to encode the internal representation to obtain a target feature, and determining the probability that the large language model outputs risky content during content generation according to the target feature.

[0094] In one or more embodiments of this specification, the target classifier is used to encode the internal representation and extract key features that are helpful in determining the risk content, and use the key features for subsequent risk probability estimation. The target classifier can be a separate machine learning model, such as logistic regression, support vector machine, or neural network model, etc., and the appropriate model architecture can be selected according to the specific task requirements, which is not limited in the embodiments of this specification.

[0095] Exemplarily, the internal representation obtained in step S304 can be used as input, mapped to the feature space through the target classifier, and feature selection can be performed using methods such as recursive feature elimination and LASSO (a shrinkage regression model) regression to screen out features that are highly correlated with the risk category from the mapped features, that is, to obtain the target features.

[0096] In one embodiment of the present specification, the target classifier may include a linear transformation layer and a nonlinear activation function. Accordingly, the linear transformation layer may be used to perform weighted summation on the internal representation, and the target feature may be calculated based on the summation result and the bias term. Then, the nonlinear activation function may be used to convert the target feature into the probability of the large language model outputting risky content during the content generation process.

[0097] For example:

[0098] label = δ(w*emb+b) (2)

[0099] Among them, δ is a nonlinear activation function, such as the softmax function, w is the weight, b is the bias term, emb is the internal representation, the calculation result of w*emb+b is the target feature, and lable represents the probability of the large language model outputting risky content.

[0100] The probability of the current large language model outputting risky content can be calculated according to formula (2). Of course, the probability of the large language model outputting risky content after s steps can also be calculated according to formula (2). That is, the internal representation after s steps is substituted into formula (2) to predict the probability of the large language model outputting risky content after s steps. Before the risky content is exposed, it is possible to make a risk prediction on whether the content output by the large language model after several steps contains risky content, thereby improving the risk management efficiency of generative content.

[0101] In the embodiments of this specification, the output vector of the middle layer of the large language model is selected to realize the risk prediction of the generated content, rather than relying solely on the output of the last layer. By analyzing the internal representation of the middle layer, it is possible to predict earlier whether the generated content is risky and take corresponding control measures in a timely manner, thereby improving the risk prevention and control performance of the large language model in terms of timeliness and effectiveness.

[0102] The embodiments of this specification can make predictions before the large language model starts to generate potential risk content, without having to wait for the large language model to generate all word units before performing full-text detection, thus shortening the detection response time and greatly improving the risk prediction efficiency of the large language model. Moreover, by utilizing information at different levels within the large language model, the accuracy of risk prediction can be improved. In addition, it should be noted that the Probing detection method in the embodiments of this specification does not require segmentation of word units, and there is no semantic loss or performance loss, thereby further improving the accuracy of risk detection.

[0103] In one or more embodiments of the present specification, by analyzing the internal representation of the middle layer of the large language model, the risk probability of the generated content can be monitored and evaluated in real time during the content generation process. Once content that may be risky (such as involving sensitive topics, inappropriate remarks, etc.) is detected, corresponding risk control measures can be taken immediately, thereby effectively improving the risk control efficiency of the generated content.

[0104] Optionally, after obtaining the probability that the large language model outputs risky content during the content generation process, the probability value can be compared with the preset risk probability to determine whether there is risky content in the model-generated content, thereby further determining the corresponding risk control measures. Among them, the specific value of the preset risk probability can be set according to the specific application scenario and business needs. For example, historical data can be referenced to determine the distribution of risky content, and a reasonable risk probability threshold can be set based on the analysis results.

[0105] Optionally, when the probability that the large language model outputs risky content during the content generation process is greater than or equal to a preset risk probability, controlling the large language model to stop generating content can effectively curb the occurrence of risk exposure.

[0106] The embodiments of this specification can monitor the generated content of the large language model in real time, and take immediate measures when risks are detected, such as interrupting content generation, issuing warnings, or modifying the generation strategy of the large language model. Among them, modifying the generation strategy of the large language model may include regenerating text content until the probability value of the risk content included in the part of the text content is less than the preset risk probability; it may also include replacing the detected risk content with synonyms; it may also include using specific prompts or constraints to guide the large language model to generate safer text content, which is not described or limited in detail in the embodiments of this specification. While being able to improve the robustness and security of the large language model, it can also enhance the user experience.

[0107] When the probability that the large language model outputs risky content during the content generation process is less than the preset risk probability, it does not affect the content generation of the large language model. At this time, new words continue to be generated through the large language model, and the new words are used to update the current word sequence, such as adding the generated new words to the current word sequence, and serving as the basis for the next iteration. Further, the internal representation of the intermediate network layer of the large language model in the content generation process is updated according to the updated current word sequence, so as to determine the probability of the large language model outputting risky content when generating the next word based on the updated internal representation using the target classifier. This process refers to the detailed description of steps S204 to S206 in one embodiment of this specification, or to the detailed description of steps S304 to S306 in another embodiment of this specification, and will not be repeated here.

[0108] In the embodiments of this specification, the output vector of the middle layer of the large language model is used to predict the risk of generative content, which can not only realize real-time risk assessment, but also enhance the security management capability of generated content, providing users with a healthier and more positive information environment.

[0109] In one or more embodiments of the present specification, before using a target classifier to determine the probability that a large language model outputs risky content during the content generation process, the target classifier can be pre-trained so that the target classifier can accurately identify whether the generated content contains risky content based on the selected intermediate layer representation, which can not only improve the quality of the generated content, but also ensure the security of the generated content.

[0110] Optionally, a training data set containing risk labels is obtained, and each sample in the training data set includes sample data and a corresponding risk label. The risk label indicates whether there is risk content in the corresponding generated content, and can be used to supervise the learning of the target classifier. The risk label can be provided by a manual auditor or can be derived from an existing high-quality database. The embodiments of this specification do not limit the source of the risk label.

[0111] First, the trained large language model is used to generate an internal representation of each sample data in the training data set, and the internal representation is input into a preset neural network to obtain the risk prediction probability of each sample data. The preset neural network in the embodiment of this specification can be stacked by a multi-layer network structure, such as a multi-layer hidden layer network.

[0112] Among them, the sample data to be analyzed can be first converted into a format acceptable to the large language model, such as word segmentation, embedding, etc., and then the processed text is input into the pre-trained large language model, and the output of the intermediate layer selected by the large language model is captured. In addition, the internal representation obtained can be further processed as needed, such as taking the average, maximum pooling, etc., to obtain a vector representation of a fixed length. For each sample data, the internal representation corresponding to each sample data can be passed to each layer of the preset neural network, and the activation function is applied in turn until the risk prediction probability of each sample data is finally output.

[0113] Then, the loss function is calculated based on the risk prediction probability of each sample data and the corresponding risk label, and the model parameters in the preset neural network are iteratively adjusted according to the calculated loss value. The training goal is to minimize the difference between the risk probability output by the target classifier and the corresponding risk label, such as using the cross entropy loss function, mean square error, etc. to measure the difference between the risk probability and the corresponding risk label.

[0114] For example, the model parameters in the preset neural network can be adjusted according to the calculated loss value through the back propagation algorithm until the preset iteration termination condition is met, and the target classifier can be obtained. The preset iteration termination condition includes reaching the maximum number of iterations, loss function convergence, etc. In addition, if a new risk type is detected in the generated content, the new risk content can be marked and used for further training.

[0115] In the embodiments of this specification, by continuously monitoring the intermediate layer output of the large language model, the model parameters can be adjusted in time during the training phase to reduce the probability of generating bad information, thereby improving the overall security and reliability of the large language model.

[0116] See also Figure 5 , is a schematic diagram of the structure of a generative content risk prediction device provided in an embodiment of this specification. Figure 5 As shown, the generative content risk prediction device 1 can be implemented as all or part of an electronic device through software, hardware or a combination of both. According to some embodiments, the generative content risk prediction device 1 includes a risk problem acquisition module 11, an internal representation determination module 12 and a risk content prediction module 13, specifically including:

[0117] The risk problem acquisition module 11 is used to acquire the risk problem text;

[0118] An internal representation determination module 12, used for inputting the risk problem text into a large language model for content generation, and extracting the internal representation of the intermediate network layer of the large language model during the content generation process;

[0119] The risk content prediction module 13 is used to determine the probability of the large language model outputting risky content during the content generation process based on the internal representation and using a target classifier.

[0120] Optionally, the intermediate network level includes at least one hidden layer in the middle position of the large language model network structure, and the generated content of the large language model includes a word sequence; when the internal representation determination module 12 extracts the internal representation of the intermediate network level of the large language model in the content generation process, it is specifically used to:

[0121] Extracting output vectors of each hidden layer during the content generation process;

[0122] The output vectors are concatenated to obtain the internal representation.

[0123] Optionally, when the internal representation determination module 12 extracts the output vectors of each hidden layer in the content generation process, it is specifically used to:

[0124] The output vector of each hidden layer when processing the risk question text and not generating word-units is extracted.

[0125] Optionally, when the internal representation determination module 12 extracts the output vectors of each hidden layer in the content generation process, it is specifically used to:

[0126] Extracting the output vector corresponding to the target word in the currently generated word sequence from each of the hidden layers;

[0127] The target word is the last word in the currently generated word sequence.

[0128] Optionally, when the risk content prediction module 13 determines the probability of the large language model outputting risk content in the content generation process based on the internal representation and using the target classifier, it is specifically used to:

[0129] Encoding the internal representation using the target classifier to obtain target features;

[0130] The probability that the large language model outputs risky content during content generation is determined according to the target feature.

[0131] Optionally, the target classifier includes a linear transformation layer; when the risk content prediction module 13 encodes the internal representation using the target classifier to obtain the target feature, it is specifically used to:

[0132] The linear transformation layer is used to perform weighted summation on the internal representation, and the target feature is calculated based on the summation result and the bias term.

[0133] Optionally, the target classifier further includes a nonlinear activation function; when the risk content prediction module 13 determines the probability that the large language model outputs risk content during the content generation process according to the target feature, it is specifically used to:

[0134] The target feature is converted into the probability of the large language model outputting risky content during the content generation process using the nonlinear activation function.

[0135] Optionally, the generative content risk prediction device 1 further includes a first risk management and control module, which is specifically used to:

[0136] When the probability that the large language model outputs risky content during the content generation process is greater than or equal to a preset risk probability, the large language model is controlled to stop generating content.

[0137] Optionally, the generative content risk prediction device 1 further includes a second risk management and control module, which is specifically used to:

[0138] When the probability that the large language model outputs risky content during the content generation process is less than a preset risk probability, generating a new word by using the large language model;

[0139] The new word-gram is used to update the current word-gram sequence and the internal representation of the intermediate network layer of the large language model during the content generation process, so as to determine the probability of the large language model outputting risky content when generating the next word-gram based on the updated internal representation and using the target classifier.

[0140] Optionally, the generative content risk prediction device 1 further includes a classifier training module, which is specifically used to:

[0141] Obtain a training dataset containing risk labels;

[0142] Generate an internal representation of each sample data in the training data set using the trained large language model;

[0143] Inputting the internal representation into a preset neural network to obtain the risk prediction probability of each sample data;

[0144] The loss function is calculated based on the risk prediction probability of each sample data and the corresponding risk label, and the model parameters in the preset neural network are iteratively adjusted according to the calculated loss value. When the iteration termination condition is met, the target classifier is obtained.

[0145] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.

[0146] The present specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figures 2 to 4 The method of the embodiment shown in the figure can be specifically executed by referring to Figures 2 to 4 The specific description of the illustrated embodiment will not be repeated here.

[0147] The present specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 2 to 4 The method of the embodiment shown in the figure can be specifically executed by referring to Figures 2 to 4 The specific description of the illustrated embodiment will not be repeated here.

[0148] The embodiments of this specification also provide Figure 6 The structural diagram of the electronic device shown in FIG. Figure 6 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned voice activity detection method.

[0149] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the executor of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0150] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0151] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.

[0152] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0153] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0154] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0155] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0156] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0158] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0159] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0160] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0161] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0162] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0163] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0164] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0165] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.

Claims

1. A generative content risk prediction method, the method comprising: Get the risk question text; Inputting the risk question text into a large language model for content generation, and extracting the internal representation of the intermediate network layer of the large language model in the content generation process; Based on the internal representation, a target classifier is used to determine the probability that the large language model outputs risky content during content generation.

2. According to the generative content risk prediction method of claim 1, the intermediate network level comprises at least one hidden layer in the middle position of the large language model network structure, and the generated content of the large language model comprises a word sequence; The extracting the internal representation of the intermediate network layer of the large language model in the content generation process includes: Extracting output vectors of each hidden layer during the content generation process; The output vectors are concatenated to obtain the internal representation.

3. The generative content risk prediction method according to claim 2, wherein the step of extracting the output vector of each hidden layer during the content generation process comprises: The output vector of each hidden layer when processing the risk question text and not generating word-units is extracted.

4. The generative content risk prediction method according to claim 2, wherein the step of extracting the output vector of each hidden layer during the content generation process comprises: Extracting the output vector corresponding to the target word in the currently generated word sequence from each of the hidden layers; The target word is the last word in the currently generated word sequence.

5. The generative content risk prediction method according to any one of claims 1 or 2, wherein the method of determining the probability of the large language model outputting risky content during content generation using a target classifier based on the internal representation comprises: Encoding the internal representation using the target classifier to obtain target features; The probability that the large language model outputs risky content during content generation is determined according to the target feature.

6. The generative content risk prediction method according to claim 5, wherein the target classifier comprises a linear transformation layer; The step of encoding the internal representation using the target classifier to obtain target features includes: The linear transformation layer is used to perform weighted summation on the internal representation, and the target feature is calculated based on the summation result and the bias term.

7. The generative content risk prediction method according to claim 5, wherein the target classifier further comprises a nonlinear activation function; The determining, according to the target feature, the probability that the large language model outputs risky content during content generation includes: The target feature is converted into the probability of the large language model outputting risky content during the content generation process using the nonlinear activation function.

8. The generative content risk prediction method according to claim 1, further comprising: When the probability that the large language model outputs risky content during the content generation process is greater than or equal to a preset risk probability, the large language model is controlled to stop generating content.

9. The generative content risk prediction method according to claim 1, further comprising: When the probability that the large language model outputs risky content during the content generation process is less than a preset risk probability, generating a new word by using the large language model; The new word-gram is used to update the current word-gram sequence and the internal representation of the intermediate network layer of the large language model during the content generation process, so as to determine the probability of the large language model outputting risky content when generating the next word-gram based on the updated internal representation and using the target classifier.

10. The generative content risk prediction method according to claim 1, further comprising: Obtain a training dataset containing risk labels; Generate an internal representation of each sample data in the training data set using the trained large language model; Inputting the internal representation into a preset neural network to obtain the risk prediction probability of each sample data; The loss function is calculated based on the risk prediction probability of each sample data and the corresponding risk label, and the model parameters in the preset neural network are iteratively adjusted according to the calculated loss value. When the iteration termination condition is met, the target classifier is obtained.

11. A generative content risk prediction device, the device comprising: Risk problem acquisition module, used to obtain risk problem text; An internal representation determination module, used for inputting the risk question text into a large language model for content generation, and extracting the internal representation of the intermediate network layer of the large language model during the content generation process; A risky content prediction module is used to determine the probability of the large language model outputting risky content during the content generation process based on the internal representation and using a target classifier.

12. A storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.

13. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 10.

14. A computer program product having at least one instruction stored thereon, wherein when the at least one instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Cited By

  • AI large model content generation security detection method and system

    CN120915984A