Honeypot-based infringement detection method, device and equipment and storage medium
By generating honeypot scripts in the script workshop system and analyzing the responses from external large models, the problem of the script workshop system's lack of control over the processing of external large model data is solved, and the detection and prevention of infringement by external large models are realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QIYI CENTURY SCI & TECH CO LTD
- Filing Date
- 2026-02-04
- Publication Date
- 2026-06-05
AI Technical Summary
The script workshop system lacks effective control over the underlying data processing logic of external AI models, which may lead to the unauthorized inclusion of high-value script texts into the training corpus by external models, resulting in infringement. It is difficult to effectively detect whether external models are infringing.
By obtaining the target character names and preset honeypot character information from the original script, a complete honeypot script is generated and fed to an external large model. According to preset rules, query information is generated, and the response results of the external large model are analyzed to detect whether there is any infringement.
It has enabled effective detection of infringement by external large-scale models, ensuring that script texts are not used without authorization, and improving the data processing and control capabilities of the script workshop system.
Smart Images

Figure CN122153849A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a honeypot-based infringement detection method, apparatus, device, and storage medium. Background Technology
[0002] With the development of technology, the application of artificial intelligence large model technology in the field of script creation is becoming more and more in-depth. The script workshop system has also achieved intelligent upgrades in functions such as script creation, evaluation, and analysis.
[0003] In practice, the data interaction between the script workshop system and the external artificial intelligence big model (hereinafter referred to as the external big model) often needs to be transferred through the Qizhi platform. After passing through the Qizhi platform, a large amount of high-value script text resources are input into the external big model, specifically including core contents such as complete scripts, script evaluation materials, and script breakdown and analysis data.
[0004] However, because the external large-scale model is set up outside the Script Workshop system, the Script Workshop system lacks effective control over its underlying data processing logic. This creates a potential risk that the external large-scale model may include the aforementioned high-value script texts in its training corpus without authorization, which could easily lead to copyright infringement by the external large-scale model provider. Against this backdrop, how to effectively detect whether the external large-scale model has engaged in such infringement has become a critical issue that urgently needs to be addressed. Summary of the Invention
[0005] This application provides a honeypot-based infringement detection method, apparatus, device, and storage medium. The method can detect whether there is any infringement in a large external model.
[0006] Firstly, this application provides a honeypot-based infringement detection method, the method comprising: Obtain the original script, the names of the target characters in the original script, and the information of the pre-defined honeypot characters; Searching the original script yields target-related plots associated with the target character's name; Based on the target-related plot and the information about the honeypot characters, the complete script for the honeypot is determined; After feeding the external large model with the complete honeypot script, query information is generated according to preset rules; Obtain the response of the external large model to the query information; Based on the response results, detect whether the external large model infringes on any rights.
[0007] Optionally, the method further includes: The original script is analyzed to determine the target script style and target script type. The target dimensions required to determine the honeypot character information corresponding to the target script type; The honeypot character information is generated based on the target dimension and the target script style.
[0008] Optionally, after obtaining the complete honeypot script, the method further includes: Check whether the complete script for the honeypot includes information about the honeypot characters; If the complete honeypot script does not include the honeypot character information, the complete honeypot script is regenerated based on the honeypot character information.
[0009] Optionally, detecting whether the external large model infringes on copyright based on the response result includes: Based on the complete honeypot script, determine the honeypot characteristics; Extract target features from the answer results; When the honeypot features and the target features meet preset conditions, it is determined that the external large model has infringed on copyright. When the honeypot features and the target features do not meet the preset conditions, it is determined that the external large model does not involve infringement.
[0010] Optionally, determining the complete script for the honeypot based on the target-related plot and the honeypot character information includes: Using the target-related plot as context and combining it with the honeypot character information, a honeypot script outline is generated; Using the target-related plot as context, and combining it with the honeypot script outline, a complete honeypot script is generated.
[0011] Optionally, query information can be generated according to preset rules, including: Obtain target association information related to the honeypot character information in the complete script of the honeypot; Based on the target association information, query information is generated.
[0012] Optionally, the method further includes: Obtain character information for each character in the original script; Based on the character information of each character, determine the character rating for each character; Based on each character's rating, the target character is identified from all characters, and the target character's information is determined as the target character's name.
[0013] Secondly, this application provides a honeypot-based infringement detection device, the device comprising: The first acquisition unit is used to acquire the original script, the names of the target characters in the original script, and the preset honeypot character information; The retrieval unit is used to search the original script to obtain target-related plots related to the name of the target character. The first determining unit is used to determine the complete script of the honeypot based on the target-related plot and the honeypot character information; The first generation unit is used to generate query information according to preset rules after feeding the external large model with the complete script of the honeypot. The second acquisition unit is used to acquire the external large model's response to the query information; The first detection unit is used to detect whether the external large model has any infringement behavior based on the response result.
[0014] Optionally, the apparatus further includes a second generating unit, the second generating unit being configured to: The original script is analyzed to determine the target script style and target script type. The target dimensions required to determine the honeypot character information corresponding to the target script type; The honeypot character information is generated based on the target dimension and the target script style.
[0015] Optionally, after obtaining the complete honeypot script, the device further includes a second detection unit, the second detection unit being used for: Check whether the complete script for the honeypot includes information about the honeypot characters; If the complete honeypot script does not include the honeypot character information, the complete honeypot script is regenerated based on the honeypot character information.
[0016] Optionally, the first detection unit is used for: Based on the complete honeypot script, determine the honeypot characteristics; Extract target features from the answer results; When the honeypot features and the target features meet preset conditions, it is determined that the external large model has infringed on copyright. When the honeypot features and the target features do not meet the preset conditions, it is determined that the external large model does not involve infringement.
[0017] Optionally, the first determining unit is configured to: Using the target-related plot as context and combining it with the honeypot character information, a honeypot script outline is generated; Using the target-related plot as context, and combining it with the honeypot script outline, a complete honeypot script is generated.
[0018] Optionally, the first generating unit is used for: Obtain target association information related to the honeypot character information in the complete script of the honeypot; Based on the target association information, query information is generated.
[0019] Optionally, the apparatus further includes a second determining unit, the second determining unit being configured to: Obtain character information for each character in the original script; Based on the character information of each character, determine the character rating for each character; Based on each character's rating, the target character is identified from all characters, and the target character's information is determined as the target character's name.
[0020] Thirdly, this application provides a honeypot-based infringement detection device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the processor is configured to: Obtain the original script, the names of the target characters in the original script, and the information of the pre-defined honeypot characters; Searching the original script yields target-related plots associated with the target character's name; Based on the target-related plot and the information about the honeypot characters, the complete script for the honeypot is determined; After feeding the external large model with the complete honeypot script, query information is generated according to preset rules; Obtain the response of the external large model to the query information; Based on the response results, detect whether the external large model infringes on any rights.
[0021] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described honeypot-based infringement detection method.
[0022] Compared with the prior art, the technical solution provided in this application has the following advantages: This application first obtains the original script, the target character names in the original script, and preset honeypot character information; then, it retrieves and extracts target-related plots related to the target character names from the original script, and combines the target-related plots with the honeypot character information to generate a complete honeypot script; after feeding this complete honeypot script to an external large model, it generates query information according to preset rules; after obtaining the external large model's response to the query information, it detects whether the external large model has engaged in infringement based on the response. In summary, this application can effectively detect infringement by external large models. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0026] Figure 1 A schematic flowchart of a honeypot-based infringement detection method provided in an embodiment of this application; Figure 2 A schematic diagram illustrating a honeypot-based infringement detection method provided in an embodiment of this application; Figure 3 A flowchart illustrating a method for determining honeypot character information provided in an embodiment of this application; Figure 4 A flowchart illustrating a method for determining infringement as provided in an embodiment of this application; Figure 5 A flowchart illustrating another method for determining the name of a target person provided in this application embodiment; Figure 6 A schematic flowchart of a honeypot-based infringement detection device provided in an embodiment of this application; Figure 7 This is a schematic diagram of a honeypot-based infringement detection device provided in an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] The following disclosure provides numerous different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of the invention. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0029] With the development of technology, the application of AI large-scale modeling technology in scriptwriting is becoming increasingly in-depth. The Script Workshop system has also achieved intelligent upgrades in scriptwriting, evaluation, and analysis functions. In practice, data interaction between the Script Workshop system and external AI large-scale models often requires the Qizhi platform as an intermediary. After passing through the Qizhi platform, a large amount of high-value script text resources are input into the external large-scale model, specifically including complete scripts, script evaluation materials, and script analysis data. However, because the external large-scale model is set up outside the Script Workshop system, the Script Workshop system lacks effective control over the underlying data processing logic of the external large-scale model. This creates a potential risk that the external large-scale model may include the aforementioned high-value script texts into its training corpus without authorization, which could easily lead to infringement by the external large-scale model provider. Against this backdrop, how to effectively detect whether large external models exhibit the aforementioned infringements has become a critical issue that urgently needs to be addressed. In summary, this application provides a honeypot-based infringement detection method that can detect whether a large external model exhibits infringing behavior, such as... Figure 1 As shown, the specific steps include: Step 101: Obtain the original script, the names of the target characters in the original script, and the information of the preset honeypot characters.
[0030] In this context, the original script refers to a completed script. In practical application, each generated script is simultaneously stored in the script database. Based on this database, unused scripts can be selected as the original script, or the most recently stored script can be extracted. Other acquisition methods are also possible, and this application does not specifically limit their application.
[0031] The target character's name is the name of a pre-designated character in the original script. The methods for determining the target character may include, but are not limited to, the following: identifying the protagonist of the original script as the target character; identifying a character with distinct personality traits in the original script as the target character; or identifying a character with high frequency of appearance in the original script as the target character. This application does not limit the specific method for determining the target character.
[0032] There are two methods for configuring honeypot character information: one is manual setting by technical personnel based on practical experience; the other is automatic generation according to preset rules. For example, if the honeypot character information only needs to include the character's name, the corresponding honeypot character name can be directly determined by randomly generating a string. To balance the dual needs of avoiding external large-scale models from identifying honeypot features and ensuring accurate honeypot identification in the future, unique distinguishing identifiers need to be embedded in the honeypot character information. For example, in addition to the honeypot character's name, the honeypot character information can also include the honeypot character's unique behaviors and lines.
[0033] When honeypot character information includes multiple pieces of information, it can be generated by combining it with the target character's information. This combination significantly improves the honeypot character's concealment and facilitates accurate subsequent identification, making it a superior generation strategy. The specific combination approach can be described as follows: Honeypot character information can be generated by combining the target character's information and unique information. On one hand, based on the target character's core attributes (such as name, behavioral characteristics, and dialogue style), corresponding information for the honeypot character can be configured to make it similar to the target character in surface features, thereby reducing the probability of being identified by external large-scale models. On the other hand, to ensure the accuracy of subsequent identification, unique distinguishing identifiers need to be embedded in the honeypot character information. For example, specific execution details can be added to actions similar to the target character, and unique keywords can be added to dialogues with similar styles.
[0034] Step 102: Search the original script to obtain target-related plots associated with the target character's name.
[0035] In this step, a set of relevant plot points for the target character is found in the original script. If the plot points include multiple elements, the most crucial plot point in the set is selected as the target-related plot. Specifically, based on the target character's name, all plot segments in the original script are traversed, and plot points that meet any of the following conditions are included in the target character's relevant plot set: the plot contains direct actions of the target character, and these actions are highly related to the target character's core tasks; the plot development revolves around the target character's role and plays a crucial role in shaping the target character's image; the plot contains key interactions between the target character and other core characters, and these interactions affect the main plot direction. In practice, the original script and the target character's name can be input into an internal large-scale model, allowing the model to search the original script based on the target character's name to obtain target-related plot points.
[0036] Among them, the internal large model is the artificial intelligence large model built into the script workshop system. It is a dedicated large model with the ability to analyze script content, generate natural language, sort out plot logic, and extract features.
[0037] Step 103: Determine the complete script for the honeypot based on the target-related plot and honeypot character information.
[0038] In this step, the aforementioned internal large model can be used to generate a complete honeypot script based on the target-related plot and honeypot character information. Alternatively, the target-related plot can be used as context, combined with honeypot character information and preset target prompts, to generate a honeypot script outline; the target-related plot can also be used as context, combined with the honeypot script outline, to generate a complete honeypot script.
[0039] Step 104: After feeding the external large model with the complete honeypot script, generate query information according to preset rules.
[0040] The preset rules can be as follows: Extract the target character's name and the honeypot character's name from the complete honeypot script, and fill these two names into the preset query templates to generate query information. For example, if the query template is "What happened between character A and character B?", then fill in the target character's name in the "character A" position and the honeypot character's name in the "character B" position to generate the corresponding query information. Other rules can also be used to generate query information; this is not limited to these rules.
[0041] Step 105: Obtain the external large model's response to the query information.
[0042] In this step, after the query information is input into the external large model, the external large model will generate the corresponding answer based on the query information.
[0043] The external large model is an AI large model that can be called by the script workshop system. This model is independently developed by a professional company and its data processing capabilities are superior to the internal large model built into the system. Therefore, it is necessary to call this external large model to carry out data processing work in relevant processes.
[0044] Step 106: Based on the response results, detect whether the external large model has any infringement behavior.
[0045] In this step, the script workshop system can detect whether there are preset keywords in the answer results and whether the semantic similarity between the answer results and the honeypot features is greater than a preset value. When there are preset keywords in the answer results and the semantic similarity between the answer results and the honeypot features is greater than the preset value, it is determined that the external large model has infringing behavior; otherwise, it is determined that the external large model has not infringed behavior.
[0046] Since machine detection has a certain degree of error, the complete honeypot script, query information, answer information, detection results and other data can be sent to relevant personnel so that they can use this data to determine whether the external large model really has infringing behavior.
[0047] This application first obtains the original script, the names of target characters within the original script, and pre-defined honeypot character information. Then, it retrieves and extracts target-related plot elements from the original script that relate to the target character names. Combining these target-related plot elements with the honeypot character information, it generates a complete honeypot script. After feeding this complete honeypot script to an external large-scale model, it generates query information according to pre-defined rules. After obtaining the external large-scale model's response to the query information, it detects whether the external large-scale model has engaged in infringement. In summary, this application can effectively detect infringement by external large-scale models.
[0048] Based on the above method, this application provides an interaction method, which is specifically as follows: Figure 2 As shown, the Script Workshop system includes a Script Workshop, a Honeypot Generation Service, a Knowledge Base, a Large Model Client (the aforementioned internal large model), the Qizhi Platform, a Database, OSS (Object Storage Service), and a Detection Service. The Script Workshop is a professional tool / platform focusing on the entire scriptwriting process; the Honeypot Generation Service is a professional tool / platform focusing on the honeypot script generation process; the Knowledge Base stores original scripts; the Large Model Client generates honeypot script outlines and content; the Qizhi Platform is a data transfer platform; and the Database stores the metadata of the honeypot scripts, enabling subsequent retrieval of honeypot scripts in OSS based on the metadata. The metadata is... Specifically, in the entire interaction process of the Scriptworks system, after the Scriptworks initiates a script upload request, the honeypot generation service calls the dataset API to upload the original script. Once the honeypot generation service successfully stores the original script, it uses it as a knowledge base (context) to request the large model client to generate a honeypot script outline. After the large model client returns the generated honeypot script outline, it continues to use the original script as a knowledge base and requests the large model client to generate a new honeypot script based on the outline. This honeypot script includes multiple chapters. The honeypot generation service then obtains the honeypot script's metadata and stores it in the database. After successful storage in the database, the honeypot script is stored in OSS. When OSS successfully stores the honeypot script, the honeypot generation service returns a honeypot script generation completion notification to the Scriptworks.
[0049] Subsequently, the Scriptworkshop initiates a honeypot feeding task to the Qizhi platform, prompting the platform to request honeypot scripts to be fed from the honeypot generation service. Upon receiving the request, the honeypot generation service queries the honeypot script metadata in its database, downloads the script from OSS based on the metadata, and then sends it to the Qizhi platform. After receiving the honeypot script, the Qizhi platform calls an external large model, delivers the script to it, receives a successful feeding notification from the external model, and sends a successful feeding notification to the Scriptworkshop. Upon receiving the successful feeding notification, the Scriptworkshop sends a detection recall task to the detection service. Upon receiving the detection recall task, the detection service generates a targeted query and calls the Qizhi platform's API to send the targeted query to the platform. Upon receiving the targeted query, the Qizhi platform initiates a query to the external large model based on the query, receives the query result from the external large model, and sends the result to the detection service. The detection service receives the query results, sends the comparison results and honeypot features to the database, retrieves the feature data from the database, and returns a detection report to the script workshop.
[0050] In this embodiment of the application, a character template can be preset, and the information required by the preset character template can be generated according to the preset character template. For example, when the information required by the preset character template is the honeypot character name, honeypot character behavior and honeypot character lines, the internal large model is called to generate this information, and the generated information is used as the honeypot character information.
[0051] Because different scripts correspond to different genres, the required information for the "honeypot" characters may also differ. For example, when a script is a traditional Chinese style script, its dialogue has strong stylistic and unique characteristics, such as containing classical Chinese sentences and ancient-style titles (e.g., "Gongzi," "Qing," "Gezhu"). Therefore, the honeypot character information for this type of script can include the honeypot character's name, behavior, and dialogue. When a script is a campus-themed script, its scenes have strong unique recognizability, such as classrooms (e.g., Class 2 of Senior Three), the third floor of the library's humanities section, the school playground's rostrum, and club activity rooms. These are all concrete and unique scenes in campus scripts. Therefore, the honeypot character information for this type of script can include the honeypot character's name, behavior, and the scene in which the honeypot character is located. Therefore, this application can also analyze the original script to determine its target script style and type; then, based on the target script type, clarify the target dimensions required for the corresponding honeypot character information; finally, combine the target dimensions and the target script style to generate suitable honeypot character information. As can be seen, the generation process of honeypot character information involved in this application combines the dual features of script type and script style. It matches the content identification features of different script types through the target dimension, and relies on script style to ensure the stylistic consistency between the honeypot character information and the original script. This makes the generated honeypot character information more closely match the characteristics of the original script, and the honeypot content more realistic. It avoids the honeypot information being judged as abnormal text by external large-scale models due to inconsistencies with the script itself, ensuring that external large-scale models can correctly capture the honeypot features. Therefore, this application provides a method for determining honeypot character information, as follows: Figure 3 As shown, the specific steps include: Step 301: Analyze the original script to determine the target script style and target script type.
[0052] Among them, the target script style refers to the overall creative style and tone of the original script, such as the target script style being ancient Chinese style, suspense style, comedy style, etc. The target script type refers to the script category based on the core characteristics of the script, such as the story background, narrative scene, core theme, etc., such as the target script type being ancient Chinese style script, campus script, workplace script, etc.
[0053] Step 302: Determine the target dimensions required for the honeypot character information corresponding to the target script type.
[0054] The target dimension refers to the dimension corresponding to the required honeypot character information. Specifically, the target dimension includes the character name, character behavior, character's scene, character dialogue, and character appearance.
[0055] In this step, different script types correspond to different honeypot character information dimensions. Thus, corresponding character information dimensions are pre-set for each script type. When the target script type is determined, the target dimension corresponding to the target script type is determined based on the target script type and the character information dimensions corresponding to each script type.
[0056] Step 303: Generate honeypot character information based on the target dimension and target script style.
[0057] In this step, information corresponding to each target dimension is generated, and the style of this information is adjusted according to the target script style to obtain the required honeypot character information. Currently, the target dimension and target script style can also be directly input into the AI model built into the script workshop system (hereinafter referred to as the internal big model), which is then used to generate honeypot character information based on the target dimension and target script style.
[0058] In this embodiment, after obtaining the complete honeypot script, the honeypot character information is used as the core verification item to check whether the complete honeypot script is an invalid script with no honeypot characters, missing honeypot character information, or incorrect character identities. This avoids invalid scripts from the generation stage, preventing subsequent detection operations from being unable to proceed due to the lack of core script elements. Specific steps include: detecting whether the complete honeypot script includes honeypot character information; when the complete honeypot script does not include honeypot character information, regenerating the complete honeypot script based on the honeypot character information.
[0059] In this step, natural language processing is performed on the complete honeypot script to extract character information for all characters included in the script. The extracted character information is then matched against preset honeypot character information to check if the complete honeypot script contains such information. If the match passes, the complete honeypot script is considered valid and can be directly used for subsequent honeypot inducements; if the match fails, a regeneration mechanism is triggered to regenerate the complete honeypot script based on the new character information.
[0060] In this embodiment, the honeypot feature in the honeypot script is used as an anchor point. If the honeypot feature appears in the model's response, it can be definitively proven that its training data contains the script, thereby determining that the external large model has infringed upon copyright and improving the accuracy of the determination. Therefore, this embodiment provides a method for determining infringement, as follows: Figure 4 As shown, the specific steps include: Step 401: Determine the honeypot characteristics based on the complete honeypot script.
[0061] Among them, honeypot features are unique, non-public, and easily identifiable details that are exclusive to the complete honeypot script. For example, a honeypot feature is a fictional place name "Luoyun City" in the honeypot script, or a specific code "Blue Tulip".
[0062] In this step, the complete honeypot script is analyzed in depth to extract its honeypot features.
[0063] Step 402: Extract target features from the answer results.
[0064] In this step, after obtaining the response results returned by the external large model, natural language processing is performed on the response results to extract the associated information contained therein and obtain the target features.
[0065] Step 403: When the honeypot features and the target features meet the preset conditions, it is determined that the external large model has infringing behavior; when the honeypot features and the target features do not meet the preset conditions, it is determined that the external large model has not infringing behavior.
[0066] In this step, the extracted target features are compared with preset honeypot features using similarity calculation or logical matching. When the comparison results meet preset conditions (such as similarity exceeding a threshold, high overlap of key plot points, or the appearance of fictional details unique to the script), the external large model is deemed to have infringed on copyrights, meaning the model illegally ingested the complete script of the honeypot during training. When the comparison results do not meet preset conditions (such as vague answers, errors in key details, or failure to mention fictional elements), the external large model is deemed not to have infringed on copyrights.
[0067] In this embodiment, a honeypot script outline can be generated first based on the target-related plot and honeypot character information. Then, using the target-related plot as background, a complete honeypot script can be generated based on the honeypot script outline. Specific steps include: using the target-related plot as context and combining it with the honeypot character information to generate a honeypot script outline; and using the target-related plot as context and combining it with the honeypot script outline to generate a complete honeypot script.
[0068] In this step, the target-related plot is used as context, and the honeypot character information and preset first prompt words are input into the internal large model to generate the honeypot script outline. Then, the target-related plot is used as context, and the honeypot script outline and preset second prompt words are used to generate the complete honeypot script.
[0069] After obtaining the honeypot script outline, it can be sent to relevant technical personnel for review to check for logical loopholes and other issues. If such issues are found, the technical personnel can make targeted modifications to the honeypot script outline according to actual needs, thereby ensuring the rationality and logical correctness of the content of the subsequently generated complete honeypot script.
[0070] The first prompt word instructs the internal large-scale model to generate a honeypot script outline. This first prompt word can be a preset prompt word, for example, it could be expressed as "Generating a honeypot script outline based on the target-related plot and honeypot character information." Alternatively, the first prompt word can be generated based on certain rules, such as the target-related plot, target character names, and honeypot character information. Specifically, if the first prompt word uses the template "Character A and Character B meet at location C and some events occur," then Character A can be set as the target character name, and Character B as the honeypot character name. The rule for determining location C is: check if the honeypot character information contains location dimension information; if it does, then that location information is designated as location C; otherwise, a location from the target-related plot is selected as location C. The second prompt word instructs the internal large-scale model to generate the complete honeypot script. For example, the second prompt word could be expressed as "Generating a complete honeypot script based on the target-related plot and honeypot script outline."
[0071] In this embodiment, the query information can be information set by a technician based on the complete honeypot script and actual needs, or it can be generated based on the honeypot character name and a preset template. For example, the preset template is "What happened to character A?", with character A set as the honeypot character name. In practice, to make the generated query information more accurate and to obtain more explicit evidence of infringement, this application can also obtain related information about the honeypot character from the complete honeypot script and generate query information based on this related information. The specific steps include: obtaining target related information about the honeypot character from the complete honeypot script; and generating query information based on the target related information.
[0072] In this step, after obtaining the complete script for the honeypot, the core filtering criteria are used to locate all plot segments, dialogues, and behavioral descriptions involving the honeypot character, eliminating redundant information unrelated to the honeypot character. From the filtered script content, target-related information is extracted. This target-related information includes interaction-related information, content-related information, behavior-related information, and other related information. Interaction-related information includes the form and scenario of interaction between the honeypot character and the target character; content-related information includes the core communication topics, agreements reached, and disagreements between the two parties; behavior-related information includes the specific actions and behaviors of the honeypot character or the target character; and related information includes the names of mentioned third parties, the tools / platforms used, and the time / location involved.
[0073] After obtaining the target association information, query information can be generated based on this target association information. For example, the target association information can be filled into a preset query template to obtain query information, or the target association information can be input into an internal large model to obtain query information.
[0074] In addition, to fully detect whether the external large model has engaged in infringement and to obtain clear evidence of infringement when it is confirmed, multiple sets of progressively layered query information can be generated based on the target-related information. Through layer-by-layer in-depth question design, the extent to which the external large model understands the honeypot script content is gradually verified. The specific steps are as follows: set up query templates corresponding to the basic fact layer, plot development layer, and detail verification layer as needed; then, find the required information in the target-related information according to the query template, and fill in this information into the query template to generate query information.
[0075] For example, at the basic factual level: where did the target and the honeypot character meet? At the plot development level: what kind of conflict erupted after they met at that location? At the detailed verification level: how did the target and the honeypot character ultimately resolve this conflict? In this embodiment, character information for all characters is first extracted from the original script. Then, an objective score for each character is calculated based on a quantitative scoring system. Finally, the target character and its corresponding name are determined according to the scoring rules. This application uses quantitative scoring as the core basis for target character selection, sets clear selection rules, and achieves automated and accurate selection of target characters, improving the overall efficiency of script character processing. Therefore, this embodiment provides a method for determining target character names, as follows: Figure 5 As shown, the specific steps include: Step 501: Obtain character information for each character in the original script.
[0076] The character information comprises comprehensive feature information for each character in the script, including basic identity information, personality traits, plot relevance information, and character arc information. Basic identity information includes four fundamental elements: character name, gender, age, and social status. Personality traits include detailed personality descriptions. Plot relevance information includes the proportion of main and subplot scenes for each character. Character arc information indicates whether the character exhibits clear growth and development.
[0077] In this step, an internal large model can be used to identify the original script and the full-dimensional feature information corresponding to each character in the script, that is, the character information of each character.
[0078] Step 502: Determine the character score for each character based on their character information.
[0079] In this step, for each person, we obtain the dimensional information corresponding to each dimension, score the dimensional information of each dimension, obtain the dimensional score of each dimension, multiply the weight of each dimension by the dimensional score to obtain the final score of each dimension, and add up the final scores of each dimension to obtain the person's score.
[0080] The scoring rules for each dimension are as follows: For the basic identity information dimension, since it includes four basic pieces of information, the corresponding dimension score can be determined based on the missing information. For example, full marks (100 points) are awarded for not missing any information, and a certain number of points are deducted for each missing item, such as 25 points. Then, the dimension score corresponding to the basic identity information dimension is multiplied by the corresponding first weight to obtain the final score for that dimension.
[0081] For the personality trait information dimension, the corresponding dimension score is determined based on the specific personality descriptions. For example, if there are three or more specific personality descriptions, the corresponding dimension score is set to the maximum score, i.e., 100 points. For each missing specific personality description, a certain number of points are deducted, for example, 20 points. Then, the dimension score corresponding to the personality trait information dimension is multiplied by the corresponding second weight to obtain the final score for that dimension.
[0082] For the plot-related information dimension, it is determined whether the original script only includes the main plot. If it only includes the main plot, the percentage of the main plot's screen time is multiplied by the maximum score to obtain the dimension score for the plot-related information dimension. If it includes the main plot and at least one subplot, the weights corresponding to the main plot and subplots are determined based on the number of subplots. For example, when there is one subplot, the weight corresponding to the main plot is set to 0.6, and the weight corresponding to the subplot is set to 0.4; when there are two subplots, the weight corresponding to the main plot is set to 0.5, and the weight corresponding to each subplot is set to 0.25. Then, the percentage of the main plot's screen time is multiplied by its corresponding weight to obtain the first percentage, and the percentage of the subplot's screen time is multiplied by its corresponding weight to obtain the second percentage. The first and second percentages are added together to obtain the total percentage, which is then multiplied by the maximum score to obtain the dimension score for the plot-related information dimension. Finally, the dimension score for the plot-related information dimension is multiplied by its corresponding third weight to obtain the final score for that dimension.
[0083] In the above process, the weight of the main plot can be determined based on its proportion in the total plot of the original script; and the weight of the subplot can be determined based on its proportion in the total plot of the original script.
[0084] For the character arc information dimension, full marks are awarded if there is a clear growth change, and 0 marks are awarded if there is no growth change. Then, the dimension score corresponding to the character arc information dimension is multiplied by the corresponding fourth weight to obtain the final score for that dimension.
[0085] Since the plot-related information dimension has the greatest impact, it has the highest weight. The personality trait information dimension has the next greatest impact, therefore, it has the second highest weight. The first weight can be greater than or less than the fourth weight; this is not limited in this case.
[0086] Step 503: Based on the character rating of each character, identify the target character among all characters and set the target character's name as the target character name.
[0087] In this step, a rating threshold can be set, and characters with ratings no lower than the threshold can be identified as target characters. Alternatively, a preset set of characters with the highest ratings can be identified and designated as target characters. Finally, the names of these characters are designated as the target character names.
[0088] like Figure 6 As shown, this application provides a honeypot-based infringement detection device, which corresponds to the method embodiment and specifically includes: The first acquisition unit 601 is used to acquire the original script, the names of the target characters in the original script, and the preset honeypot character information; The retrieval unit 602 is used to search the original script to obtain target-related plots related to the name of the target character. The first determining unit 603 is used to determine the complete script of the honeypot based on the target-related plot and the honeypot character information; The first generation unit 604 is used to generate query information according to preset rules after feeding the external large model with the honeypot complete script. The second acquisition unit 605 is used to acquire the external large model's response to the query information; The first detection unit 606 is used to detect whether the external large model has any infringement behavior based on the answer result.
[0089] Optionally, the apparatus further includes a second generating unit 607, the second generating unit 607 being used for: The original script is analyzed to determine the target script style and target script type. The target dimensions required to determine the honeypot character information corresponding to the target script type; Based on the target dimension and the target script style, the honeypot character information is generated. Optionally, after obtaining the complete honeypot script, the device further includes a second detection unit 608, which is used for: Check whether the complete script for the honeypot includes information about the honeypot characters; If the complete honeypot script does not include the honeypot character information, the complete honeypot script is regenerated based on the honeypot character information.
[0090] Optionally, the first detection unit 606 is used for: Based on the complete honeypot script, determine the honeypot characteristics; Extract target features from the answer results; When the honeypot features and the target features meet preset conditions, it is determined that the external large model has infringed on copyright. When the honeypot features and the target features do not meet the preset conditions, it is determined that the external large model does not involve infringement.
[0091] Optionally, the first determining unit 603 is configured to: Using the target-related plot as context and combining it with the honeypot character information, a honeypot script outline is generated; Using the target-related plot as context, and combining it with the honeypot script outline, a complete honeypot script is generated.
[0092] Optionally, the first generating unit 604 is used for: Obtain target association information related to the honeypot character information in the complete script of the honeypot; Based on the target association information, query information is generated.
[0093] Optionally, the device further includes a second determining unit 609, the second determining unit 609 being configured to: Obtain character information for each character in the original script; Based on the character information of each character, determine the character rating for each character; Based on each character's rating, the target character is identified from all characters, and the target character's information is determined as the target character's name.
[0094] like Figure 7 As shown in the figure, this application provides a honeypot-based infringement detection device, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704. The processor 701, communication interface 702, and memory 703 communicate with each other via the communication bus 704. Memory 703 is used to store computer programs; In one embodiment of this application, when the processor 701 executes a program stored in the memory 703, it implements the honeypot-based infringement detection method provided in any of the foregoing method embodiments, including: Obtain the original script, the names of the target characters in the original script, and the information of the pre-defined honeypot characters; Searching the original script yields target-related plots associated with the target character's name; Based on the target-related plot and the information about the honeypot characters, the complete script for the honeypot is determined; After feeding the external large model with the complete honeypot script, query information is generated according to preset rules; Obtain the response of the external large model to the query information; Based on the response results, detect whether the external large model infringes on any rights.
[0095] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps performed by the honeypot-based infringement detection method provided in any of the foregoing method embodiments.
[0096] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0098] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0099] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A honeypot-based infringement detection method, characterized in that, The method includes: Obtain the original script, the names of the target characters in the original script, and the information of the pre-defined honeypot characters; Searching the original script yields target-related plots associated with the target character's name; Based on the target-related plot and the information about the honeypot characters, the complete script for the honeypot is determined; After feeding the external large model with the complete honeypot script, query information is generated according to preset rules; Obtain the response of the external large model to the query information; Based on the response results, detect whether the external large model infringes on any rights.
2. The method according to claim 1, characterized in that, The method further includes: The original script is analyzed to determine the target script style and target script type. The target dimensions required to determine the honeypot character information corresponding to the target script type; The honeypot character information is generated based on the target dimension and the target script style.
3. The method according to claim 1, characterized in that, After obtaining the complete honeypot script, the method further includes: Check whether the complete script for the honeypot includes information about the honeypot characters; If the complete honeypot script does not include the honeypot character information, the complete honeypot script is regenerated based on the honeypot character information.
4. The method according to claim 1, characterized in that, The step of detecting whether the external large model infringes on copyright based on the response results includes: Based on the complete honeypot script, determine the honeypot characteristics; Extract target features from the answer results; When the honeypot features and the target features meet preset conditions, it is determined that the external large model has infringed on copyright. When the honeypot features and the target features do not meet the preset conditions, it is determined that the external large model does not involve infringement.
5. The method according to claim 1, characterized in that, The step of determining the complete script for the honeypot based on the target-related plot and the honeypot character information includes: Using the target-related plot as context and combining it with the honeypot character information, a honeypot script outline is generated; Using the target-related plot as context, and combining it with the honeypot script outline, a complete honeypot script is generated.
6. The method according to claim 1, characterized in that, The step of generating query information according to preset rules includes: Obtain target association information related to the honeypot character information in the complete script of the honeypot; Based on the target association information, query information is generated.
7. The method according to claim 1, characterized in that, The method further includes: Obtain character information for each character in the original script; Based on the character information of each character, determine the character rating for each character; Based on each character's rating, the target character is identified among all characters, and the target character's information is determined as the target character's name.
8. A honeypot-based infringement detection device, characterized in that, The device includes: The first acquisition unit is used to acquire the original script, the names of the target characters in the original script, and the preset honeypot character information; A retrieval unit is used to search the original script to obtain target-related plots related to the name of the target character. The first determining unit is used to determine the complete script of the honeypot based on the target-related plot and the honeypot character information; The first generation unit is used to generate query information according to preset rules after feeding the external large model with the complete script of the honeypot. The second acquisition unit is used to acquire the external large model's response to the query information; The first detection unit is used to detect whether the external large model has any infringement behavior based on the response result.
9. A honeypot-based infringement detection device, characterized in that, include: At least one communication interface; At least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; At least one memory connected to the at least one bus, wherein the processor is configured to: Obtain the original script, the names of the target characters in the original script, and the information of the pre-defined honeypot characters; Searching the original script yields target-related plots associated with the target character's name; Based on the target-related plot and the information about the honeypot characters, the complete script for the honeypot is determined; After feeding the external large model with the complete honeypot script, query information is generated according to preset rules; Obtain the response of the external large model to the query information; Based on the response results, detect whether the external large model infringes on any rights.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the honeypot-based infringement detection method according to any one of claims 1 to 7.