A blockchain-based copyright authentication precision optimization method and system

By constructing a withdrawal/modification ratio curve and semantic similarity analysis, the credibility of ownership is dynamically adjusted, solving the problem of misjudgment of credibility in copyright authentication in existing technologies. This enables multi-dimensional evaluation of user behavior and content, improving the accuracy and robustness of copyright authentication.

CN120688039BActive Publication Date: 2026-02-27SHANDONG HANTU SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510797536.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2026-02-27
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In existing technologies, blockchain-based copyright authentication methods are unable to effectively identify unstable behaviors such as repeated claims, pseudo-original content, or habitual withdrawals by users on the platform, leading to misjudgments of the credibility of copyright authentication. Furthermore, they lack semantic analysis of the correlation between test data and users' historical unstable test questions, and cannot reasonably handle easily confused content.

Method used

By acquiring user-submitted test questions and their historical interaction records, a withdrawal/modification ratio curve is constructed. Combining user behavior trends and semantic similarity, the ownership credibility is dynamically adjusted. A first adjustment factor and a second adjustment factor are introduced, and the corrected ownership credibility is calculated based on the average slope of the withdrawal ratio curve and semantic similarity.

Benefits of technology

It improves the accuracy and credibility of copyright authentication, identifies the stability of user claims and the risk of repetitive behavior in specific content categories, enhances the ability to handle complex situations such as pseudo-original content and duplicate submissions, and improves the objectivity and robustness of rights assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688039B_ABST
    Figure CN120688039B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of blockchain copyright authentication, and provides a copyright authentication precision optimization method and system based on a blockchain, which comprises the following steps: obtaining target test question data submitted by a user and registered to a blockchain platform, obtaining an initial ownership credibility determined for the target test question data, and obtaining historical interaction records of the user in the blockchain platform. The application introduces a first adjustment factor based on user behavior trends and a second adjustment factor based on semantic similarity analysis, constructs a computer mechanism for dynamically correcting the ownership credibility of test question data, and realizes the upgrade of copyright authentication from static assertion to behavior-content bidirectional evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of blockchain copyright authentication, and particularly relates to a copyright authentication precision optimization method and system based on a blockchain. BACKGROUND

[0002] With the rapid development of educational digital content, test question data as core teaching resources are frequently created, edited and distributed in various online platforms. In order to protect the rights and interests of content creators and standardize the order of content circulation, more and more teaching platforms introduce a copyright identification mechanism based on a blockchain, and store and claim the test question data submitted by users through chain registration. However, in the prior art, static rules are often used to judge the credibility of test question claims, which mainly rely on objective indicators such as submission time, content fingerprint matching or basic information integrity, ignoring the actual behavior characteristics of users in the platform and the stability performance in the content claim process, and it is difficult to effectively identify non-stable behaviors such as repeated claims, pseudo-original creation or inertia withdrawal, resulting in the risk of credibility misjudgment in copyright authentication.

[0003] In addition, the prior art lacks analysis of the relevance between the current test question data and the user's historical unstable test questions from the semantic level, and cannot evaluate the claim rationality from the content similarity angle, resulting in that part of the easily confused content cannot be reasonably processed during the identification. Therefore, a dynamic evaluation mechanism combining user historical behavior trends and content semantic features is urgently needed to improve the precision and fairness of test question copyright authentication. SUMMARY

[0004] The purpose of the present application is to provide a copyright authentication precision optimization method and system based on a blockchain, which aims to solve the problems raised in the background art.

[0005] The present application is implemented as follows: a copyright authentication precision optimization method based on a blockchain, the method comprising:

[0006] obtaining target test question data submitted by a user and registered to a blockchain platform, and obtaining an initial ownership credibility determined for the target test question data, while obtaining historical interaction records of the user in the blockchain platform;

[0007] parsing the historical interaction records, extracting historical test question data with consistent attributes from the target test question data, calculating the withdrawal modification ratio of the target test question data and each historical test question data, and constructing a withdrawal ratio curve in chronological order;

[0008] determining whether the withdrawal modification ratio of the target test question data is located in the fitting interval of the withdrawal ratio curve, and if so, determining a first adjustment factor according to the change trend of the withdrawal ratio curve;

[0009] screening several representative test question data whose withdrawal modification ratio is higher than a preset threshold, extracting semantic vectors of the modified segments in the target test question data and each representative test question data based on semantic analysis, and sequentially calculating semantic similarity, determining a second adjustment factor based on the similarity result;

[0010] combining the first adjustment factor and the second adjustment factor, correcting the initial ownership credibility to generate a corrected ownership credibility.

[0011] As a further limitation of the technical scheme of the embodiment of the application, the step of analyzing the historical interaction records, extracting historical test question data having consistent attributes with the target test question data, calculating the withdrawal modification ratio of the target test question data and each historical test question data, and constructing a withdrawal ratio curve in chronological order comprises:

[0012] analyzing the historical interaction records, and screening historical test question data having consistent attributes with the target test question data, the consistent attributes including question type category, knowledge point label, and test question usage type;

[0013] obtaining the number of modifications, the number of withdrawals, and the total number of submissions of the target test question data and each historical test question data, and respectively calculating the ratio of the number of modifications to the total number of submissions and the ratio of the number of withdrawals to the total number of submissions as a first ratio and a second ratio, linearly weighting the first ratio and the second ratio to obtain a corresponding withdrawal modification ratio;

[0014] constructing a withdrawal ratio curve with time sequence as the horizontal axis and the withdrawal modification ratio as the vertical axis.

[0015] As a further limitation of the technical scheme of the embodiment of the application, the step of determining whether the withdrawal modification ratio of the target test question data is located in the fitting interval of the withdrawal ratio curve, and if so, determining the first adjustment factor according to the change trend of the withdrawal ratio curve comprises:

[0016] substituting the withdrawal modification ratio corresponding to the target test question data into the withdrawal ratio curve to determine whether the offset amplitude thereof in the fitting interval is less than a preset threshold;

[0017] if the offset amplitude is less than the preset threshold, calculating the average slope of the withdrawal ratio curve, and taking the average slope as the first adjustment factor.

[0018] As a further limitation of the technical scheme of the embodiment of the application, the step of screening several representative test question data whose withdrawal modification ratio is higher than a preset threshold, extracting semantic vectors of the modified segments in the target test question data and each representative test question data based on semantic analysis, and sequentially calculating semantic similarity, determining a second adjustment factor based on the similarity result comprises:

[0019] screening a preset number of representative test question data whose withdrawal modification ratio is higher than a preset threshold from the historical test question data;

[0020] extracting semantic vectors of the modified segments in the target test data and each piece of representative test data based on semantic analysis, and calculating semantic similarity between the target test data and each piece of representative test data by using cosine similarity;

[0021] calculating an average value of the preset number of semantic similarities, and taking the average value as a second adjustment factor.

[0022] As a further limitation of the technical solutions of the embodiments of the present application, after the determination of the first adjustment factor and the second adjustment factor is completed, the first adjustment factor and the second adjustment factor are combined with the initial ownership trustworthiness based on a preset ownership trustworthiness correction formula to generate a corrected ownership trustworthiness;

[0023] The ownership trustworthiness correction formula is: wherein denotes the corrected ownership trustworthiness, denotes the initial ownership trustworthiness, denotes the first adjustment factor, i.e., the average slope of the withdrawal ratio curve, denotes the adjustment coefficient corresponding to the first adjustment factor, denotes the preset number, denotes the semantic similarity of the i-th representative test data, denotes the second adjustment factor, i.e., the average value of the preset number of semantic similarities, denotes the adjustment coefficient corresponding to the second adjustment factor.

[0024] A copyright authentication precision optimization system based on a blockchain, the system comprising: a data acquisition module, a data screening module, a first adjustment factor determination module, a second adjustment factor determination module, and an ownership trustworthiness correction module, wherein:

[0025] The data acquisition module is configured to acquire target test data submitted by a user and registered on a blockchain platform, and to acquire an initial ownership trustworthiness determined for the target test data, and to acquire historical interaction records of the user on the blockchain platform.

[0026] The data screening module is configured to analyze the historical interaction records, to extract historical test data having consistent attributes with the target test data, to calculate withdrawal modification ratios of the target test data and each piece of historical test data, and to construct a withdrawal ratio curve in chronological order.

[0027] The first adjustment factor determination module is configured to determine whether the withdrawal modification ratio of the target test data is located in a fitting interval of the withdrawal ratio curve, and if so, to determine a first adjustment factor according to a change trend of the withdrawal ratio curve. ​

[0028] The second adjustment factor determination module is configured to screen a plurality of representative test question data with a withdrawal modification ratio higher than a preset threshold, extract semantic vectors of the modified segments in the target test question data and each representative test question data based on semantic analysis, and sequentially calculate semantic similarities, and determine a second adjustment factor based on the similarity results.

[0029] The ownership credibility correction module is configured to correct the initial ownership credibility based on the first adjustment factor and the second adjustment factor, and generate a corrected ownership credibility.

[0030] As a further limitation of the technical scheme of the embodiment of the present application, the data screening module specifically comprises:

[0031] The data screening unit is configured to parse the historical interaction records, and screen historical test question data with consistent attributes from the target test question data, the consistent attributes including a type category, a knowledge point label, and a test question usage type.

[0032] The ratio calculation unit is configured to obtain the number of modifications, the number of withdrawals, and the total number of submissions of the target test question data and each historical test question data, and calculate the ratio of the number of modifications to the total number of submissions and the ratio of the number of withdrawals to the total number of submissions as a first ratio and a second ratio respectively, and linearly weight the first ratio and the second ratio to obtain a corresponding withdrawal modification ratio.

[0033] The curve construction unit is configured to construct a withdrawal ratio curve with time sequence as the horizontal axis and the withdrawal modification ratio as the vertical axis.

[0034] As a further limitation of the technical scheme of the embodiment of the present application, the first adjustment factor determination module specifically comprises:

[0035] The ratio matching judgment unit is configured to substitute the withdrawal modification ratio corresponding to the target test question data into the withdrawal ratio curve, and judge whether the offset amplitude in the fitting interval is less than a preset threshold.

[0036] The average slope calculation unit is configured to calculate the average slope of the withdrawal ratio curve if the offset amplitude is less than the preset threshold, and take the average slope as the first adjustment factor.

[0037] As a further limitation of the technical scheme of the embodiment of the present application, the second adjustment factor determination module specifically comprises:

[0038] The representative data extraction unit is configured to screen a preset number of representative test question data with a withdrawal modification ratio higher than a preset threshold from the historical test question data.

[0039] The semantic similarity calculation unit is configured to extract semantic vectors of the modified segments in the target test paper data and each representative test paper data based on semantic analysis, and calculate semantic similarity between the target test paper data and each representative test paper data by using cosine similarity.

[0040] The average value calculation unit is configured to calculate an average value of the preset number of semantic similarities, and take the average value as a second adjustment factor.

[0041] As a further limitation of the technical scheme of the embodiment of the present application, after the determination of the first adjustment factor and the second adjustment factor is completed, the first adjustment factor and the second adjustment factor are combined with the initial ownership trustworthiness based on a preset ownership trustworthiness correction formula to generate a corrected ownership trustworthiness.

[0042] The ownership trustworthiness correction formula is: , wherein denotes the corrected ownership trustworthiness, denotes the initial ownership trustworthiness, denotes the first adjustment factor, i.e., the average slope of the withdrawal ratio curve, denotes an adjustment coefficient corresponding to the first adjustment factor, denotes the preset number, denotes the semantic similarity of the i-th representative test paper data, denotes the second adjustment factor, i.e., the average value of the preset number of semantic similarities, denotes an adjustment coefficient corresponding to the second adjustment factor.

[0043] Compared with the prior art, the present application has the following beneficial effects:

[0044] The present application introduces a first adjustment factor based on user behavior trends and a second adjustment factor based on semantic similarity analysis, and constructs a computer mechanism for dynamically correcting the ownership trustworthiness of test paper data, realizing the upgrade of copyright authentication from static assertion to behavior-content bidirectional evaluation. Compared with the existing technology of relying only on time stamp or content fingerprint for right confirmation, the present application can identify the stability of user's assertion in a specific content category and the risk of repeated behavior, improve the processing capability in complex situations such as pseudo-original and repeated submission, and enhance the objectivity and credibility of right confirmation evaluation. On the basis of ensuring system interpretability, the present application has good adaptability and foresight, and helps to improve the overall precision and robustness of copyright authentication on the blockchain platform. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 The flowchart of the method provided by the embodiment of the present application;

[0046] ​Figure 2 A flowchart for constructing a withdrawal rate curve in the method provided by the embodiment of the present application;

[0047] Figure 3 A flowchart for determining a first adjustment factor in the method provided by the embodiment of the present application;

[0048] Figure 4 A flowchart for determining a second adjustment factor in the method provided by the embodiment of the present application;

[0049] Figure 5 An application architecture diagram of the system provided by the embodiment of the present application;

[0050] Figure 6 A structural block diagram of a data screening module in the system provided by the embodiment of the present application;

[0051] Figure 7 A structural block diagram of a first adjustment factor determination module in the system provided by the embodiment of the present application;

[0052] Figure 8 A structural block diagram of a second adjustment factor determination module in the system provided by the embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0054] Figure 1 A flowchart of the method provided by the embodiment of the present application is shown.

[0055] Specifically, a copyright authentication precision optimization method based on a block chain, the method specifically comprises the following steps:

[0056] Step S100, obtaining target test data submitted by a user and registered to a block chain platform, and obtaining an initial ownership credibility determined for the target test data, and simultaneously obtaining historical interaction records of the user in the block chain platform.

[0057] In the embodiments of the present application, the present application is applicable to the blockchain platform scenario related to the copyright of educational content, especially applicable to the teaching resource platform, online examination platform, teaching content creation and distribution platform, etc. with the question bank system as the core. In these platforms, users usually submit structured teaching resources such as test questions as content creators, and hope to obtain a reliable copyright claim record to avoid plagiarism, dispute claims or repeated publishing behavior. The present application builds a behavior-driven dynamic adjustment mechanism, so that the copyright authentication not only depends on static information, but also comprehensively considers the historical behavior patterns of users in the platform, thereby improving the scientificity and pertinence of the reliability judgment of the copyright.

[0058] The target test question data submitted by the user and registered to the blockchain platform refers to the teaching resource data for copyright claim created or uploaded by the user in the platform, and the content confirmation, timestamp marking and data evidence are completed through the on-chain registration process of the blockchain platform. The target test question data can be a single question content, a test question book formed by multiple questions, or a structured question group document after sorting, usually including stem, options, answers, analysis, question type, applicable grade, knowledge point label and other fields. The platform can chain the core fields or abstract information in a hash manner to ensure that the content cannot be tampered with and can be traced.

[0059] The initial ownership reliability determined for the target test question data can be obtained by existing technical means. Generally, existing platforms will generate a preliminary reliability score for users based on some objective static characteristics at the initial stage of copyright, such as the order of submission time, the integrity of the content, the fingerprint repetition degree with existing resources in the platform, the identity authentication level of the submitter, whether it is the first publication, etc. These means are mainly realized through content fingerprint comparison, metadata analysis, account-level information filtering, etc. to generate a basic ownership judgment result as an input parameter for the subsequent dynamic adjustment mechanism.

[0060] The historical interaction records of the user in the blockchain platform can be recorded and synchronized on the chain by the platform, usually obtained by combining operation behavior logs and on-chain record hashes. The interaction records at least include the test question uploading behavior, modification operation, withdrawal operation, copyright claim time point, copyright failure record, dispute history, authorization behavior and historical test question data of the user in the platform. The historical test question data includes the test question content and metadata information submitted by the user, which can specifically include structured fields such as stem, options, answers, analysis, question type classification, knowledge point label, and the copyright status, modification record and version number of the test question data. The above interaction records can be extracted by cross-verification through the data structure on the blockchain side (such as the timestamp event sequence in the block, contract execution log, etc.) and the platform database to form a time-sequenced behavior data set as the basis for subsequent behavior evaluation and factor generation.

[0061] Further, the blockchain-based copyright authentication precision optimization method further includes the following steps:

[0062] Step S200, analyze the historical interaction record, extract the historical test data with consistent attributes with the target test data, calculate the withdrawal modification ratio of the target test data and each historical test data, and construct a withdrawal ratio curve in time sequence.

[0063] Specifically, Figure 2 A flowchart for constructing the withdrawal ratio curve is shown.

[0064] Among them, the analysis of the historical interaction record, the extraction of the historical test data with consistent attributes with the target test data, the calculation of the withdrawal modification ratio of the target test data and each historical test data, and the construction of the withdrawal ratio curve in time sequence specifically include the following steps:

[0065] Step S201, analyze the historical interaction record, and filter out the historical test data with consistent attributes with the target test data, the consistent attributes including the type of the question, the knowledge point label and the test purpose type;

[0066] Step S202, obtain the modification times, withdrawal times and total submission times of the target test data and each historical test data, and calculate the ratio of the modification times to the total submission times and the ratio of the withdrawal times to the total submission times, respectively, as the first ratio and the second ratio, linearly weight the first ratio and the second ratio, and obtain the corresponding withdrawal modification ratio;

[0067] Step S203, take the time sequence as the horizontal axis and the withdrawal modification ratio as the vertical axis to construct the withdrawal ratio curve.

[0068] In the embodiment of the application, the consistent attributes can not only be attributes completely consistent with the target test data, but also can be attribute combinations with similarities in the type of the question, the knowledge point label or the test purpose type. By introducing a fuzzy matching or label nesting relationship judgment mechanism, a certain degree of semantic expandability is allowed, so that the selected historical test data has strong context relevance, thereby improving the reference value of subsequent behavior analysis and the stability of the correction result.

[0069] After obtaining the modification times, the withdrawal times and the total submission times of the target test question data and each historical test question data, the ratio of the modification times to the total submission times and the ratio of the withdrawal times to the total submission times are calculated respectively as the first ratio and the second ratio. This step can be completed by searching the historical records of the corresponding test question identifier in the platform behavior log of the user, wherein the total submission times are the cumulative number of completed submissions of a test question in its entire life cycle, and the modification times and the withdrawal times are calculated based on the editing or deleting behavior operation records of the user for the test question. Then, by linear weighting of the first ratio and the second ratio, a weighted function such as: withdrawal modification ratio = a x first ratio + b x second ratio can be used, wherein a and b are the platform preset weight coefficients. The significance of this weighting calculation step is to quantitatively model the stability of the user's advocacy for a certain type of content, providing a unified standard index for subsequent trend fitting and credibility correction. Compared with analyzing only one of the modification or withdrawal behaviors, using the combined ratio can more comprehensively reflect the user's behavior tendency and content control degree.

[0070] In constructing the withdrawal ratio curve, the time sequence is taken as the horizontal axis and the withdrawal modification ratio is taken as the vertical axis, and existing data visualization and time series modeling techniques can be used to achieve this. Specifically, by reading the withdrawal modification ratio values of the target user for the same attribute test question at different time points, a time series data point set can be constructed, and then interpolation, spline fitting or regression fitting methods can be used to smooth the data point sequence to form a trend curve. The curve construction process can call existing time series analysis tools (such as the regression model in the pandas, matplotlib, scikit-learn library of Python, or the built-in function in the database). Through this step, the behavior evolution pattern of the user in a specific content category can be more clearly revealed, providing dynamic quantitative support for behavior credibility.

[0071] Further, the blockchain-based copyright authentication accuracy optimization method further includes the following steps:

[0072] Step S300, determining whether the withdrawal modification ratio of the target test question data is located in the fitting interval of the withdrawal ratio curve, if yes, determining a first adjustment factor according to the change trend of the withdrawal ratio curve.

[0073] Specifically, Figure 3 A flowchart for determining the first adjustment factor is shown.

[0074] Wherein, determining whether the withdrawal modification ratio of the target test question data is located in the fitting interval of the withdrawal ratio curve, if yes, determining a first adjustment factor according to the change trend of the withdrawal ratio curve specifically includes the following steps:

[0075] Step S301, the withdrawal modification ratio corresponding to the target test question data is substituted into the withdrawal ratio curve to determine whether the offset amplitude in the fitting interval is less than a preset threshold value;

[0076] Step S302, if the offset amplitude is less than the preset threshold value, the average slope of the withdrawal ratio curve is calculated, and the average slope is taken as the first adjustment factor.

[0077] In the embodiments of the present application, the withdrawal modification ratio corresponding to the target test question data is substituted into the withdrawal ratio curve, and it is determined whether the offset amplitude in the fitting interval is less than a preset threshold value, in order to determine whether the behavior performance of the test question data conforms to the historical behavior trend of the user on similar test question content. The determination operation can effectively filter out abnormal or non-representative data points, and avoid unreasonable interference on the credibility evaluation result caused by isolated fluctuations or accidental behaviors. When the offset amplitude is small, it means that the behavior characteristics of the test question data are highly consistent with the user's past performance on this type of test question, which has statistical representativeness and provides a basis for further slope trend analysis.

[0078] After determining that the condition is met, the average slope of the withdrawal ratio curve is calculated, which can be realized based on the derivative estimation method of time series data. Common technical means include using first-order difference method, local linear regression, least squares fitting and other methods within a preset time window to approximate the slope of multiple data points of the ratio curve.

[0079] The main significance of taking the average slope as the first adjustment factor is to introduce the dynamic correction ability of the user behavior evolution trend to the initial ownership credibility. The average slope reflects whether the user's withdrawal modification behavior on this type of test question data shows a downward trend (behavior tends to be stable) or an upward trend (behavior fluctuation intensifies). If the average slope is negative, it means that the user's stability of the claim on this type of content is enhanced, and the credibility should be increased; if it is positive, it means that the claim behavior is still unstable, and the credibility should be conservatively corrected or reduced. Therefore, this factor has the dual functions of behavior trend quantization and correction direction guidance, which can effectively enhance the rationality and timeliness of the calculation result of the right confirmation credibility.

[0080] Further, the copyright authentication precision optimization method based on the block chain further includes the following steps:

[0081] Step S400, a plurality of representative test question data with withdrawal modification ratio higher than a preset threshold value are screened, semantic vectors of the modified fragments in the target test question data and each representative test question data are extracted based on semantic analysis, and semantic similarity is calculated in sequence, a second adjustment factor is determined based on the similarity result.

[0082] Specifically, Figure 4 A flowchart for determining the second adjustment factor is shown.

[0083] The method comprises the following steps of: screening a plurality of representative test question data with a high modification and withdrawal ratio higher than a preset threshold value, extracting semantic vectors of the modified segments in the target test question data and each of the representative test question data based on semantic analysis, and sequentially calculating semantic similarities, and determining a second adjustment factor based on the similarity results.

[0084] In step S401, a preset number of representative test question data with a high modification and withdrawal ratio higher than a preset threshold value are screened from historical test question data.

[0085] In step S402, semantic vectors of the modified segments in the target test question data and each of the representative test question data are extracted based on semantic analysis, and the cosine similarity is used to calculate the semantic similarity between the target test question data and each of the representative test question data.

[0086] In step S403, an average value of the preset number of semantic similarities is calculated, and the average value is taken as the second adjustment factor.

[0087] In the embodiment of the present application, the preset threshold value can be adjusted according to the platform's classification and statistical results of user behavior stability, and a proportion boundary value with significant behavior fluctuation is usually selected, for example, after the historical distribution of the withdrawal and modification ratio of all users is counted, the top 25% or the top 10% is taken as the threshold value. The threshold value aims to screen out sample data with unstable withdrawal and modification behavior in this type of test question attribute, so as to mine the content that the user is prone to repeatedly claim or lack of confidence in the past as a reference basis for judging the current test question behavior pattern.

[0088] The preset number of representative test question data with a high modification and withdrawal ratio higher than a preset threshold value are screened from historical test question data, in order to ensure the quality of representative behavior data samples while controlling the calculation resources and similarity calculation complexity. The preset number can be set according to the system resources and the evaluation accuracy needs, such as 5, 10, 20, etc., and it is generally recommended to be within 10 to avoid the problem of sample dilution leading to too smooth semantic similarity average and weakening the discrimination.

[0089] In the implementation process of semantic analysis, the modified segments in the target test question data and the representative test question data need to be extracted first, which usually includes the difference parts of the core fields such as the stem, the option content and the analysis statement. After extraction, the text is processed by word segmentation and coding, and existing semantic vector modeling methods such as BERT, Sentence-BERT, SimCSE, etc. are used to map the text content into fixed-dimensional semantic vectors. Then, the cosine similarity is used to calculate the angle similarity between the vectors of the target test question data and each of the representative test question data. This method has been widely used in the field of natural language processing for semantic similarity calculation, and is mature and easy to deploy, supporting implementation based on mainstream frameworks such as TensorFlow and PyTorch.

[0090] Calculating the average of the preset number of semantic similarities and using this average as the second adjustment factor is significant in quantifying the semantic association strength between the current question and historically highly volatile questions. A high average semantic similarity indicates that the current question is highly similar in content expression or structure to questions frequently modified or withdrawn by the user in the past, potentially recurring in areas of unstable claims, and its credibility should be conservatively adjusted. Conversely, a low average similarity indicates that the current question possesses a degree of independence and does not belong to the "habitual content" repeatedly claimed by the user, making its claims more credible. This indicator can assist the system in identifying potential duplication, rewriting, or pseudo-original behavior at the semantic level, enhancing the content targeting and behavioral explanatory power of the overall adjustment logic.

[0091] Furthermore, the blockchain-based copyright authentication accuracy optimization method also includes the following steps:

[0092] Step S500: After determining the first adjustment factor and the second adjustment factor, based on the preset ownership credibility correction formula, the first adjustment factor and the second adjustment factor are combined with the initial ownership credibility to generate the corrected ownership credibility.

[0093] The ownership credibility correction formula is as follows: ,in This refers to the credibility of the revised ownership. This refers to the credibility of initial ownership. This refers to the first adjustment factor, which is the average slope of the withdrawal ratio curve. This refers to the adjustment coefficient corresponding to the first adjustment factor. This refers to the preset quantity. It refers to the first Semantic similarity of representative test questions This refers to the second adjustment factor, which is the average of a preset number of semantic similarities. This refers to the adjustment coefficient corresponding to the second adjustment factor.

[0094] In this embodiment of the invention, by combining the first adjustment factor and the second adjustment factor to jointly correct the initial ownership credibility, the characteristics of user claims behavior on the target test data can be more comprehensively reflected. The first adjustment factor, based on the changing trend of the withdrawal / modification ratio curve, depicts the evolution of user claim behavior over time, and has strong time sensitivity and behavioral trend expression ability; the second adjustment factor, through semantic similarity calculation, introduces the semantic correlation between the current content and the past unstable content, enhancing the ability to identify potential duplicate claims, pseudo-original content, or templated creation behavior.

[0095] The combination not only makes the credibility correction process change from single dimension to multi-dimensional behavior-content fusion, but also makes the system have more context continuity and content insight in understanding the user's claim. The additional advantage brought by this structure is that it has the ability to learn the "inertial behavior pattern" of the user, and can actively identify the content implied by the risk factors without explicitly marking the untrustworthy content, thereby realizing more forward-looking claim credibility judgment logic. Compared with the traditional rule-based or static index-based judgment method, this method shows higher adaptability and stronger abnormality recognition ability in actual deployment, providing structural support for improving the accuracy and anti-avoidance ability of the platform right authentication system.

[0096] Further, Figure 5 The application architecture diagram of the system provided by the embodiment of the application is shown.

[0097] In another preferred embodiment provided by the application, a copyright authentication precision optimization system based on a block chain comprises:

[0098] The data acquisition module 100 is configured to acquire target test question data submitted by a user and registered on a block chain platform, and to acquire an initial ownership credibility determined for the target test question data, and to acquire historical interaction records of the user on the block chain platform.

[0099] Further, the copyright authentication precision optimization system based on the block chain further comprises:

[0100] The data screening module 200 is configured to analyze the historical interaction records, to extract historical test question data having consistent attributes with the target test question data, to calculate a withdrawal modification ratio of the target test question data and each piece of historical test question data, and to construct a withdrawal ratio curve in chronological order.

[0101] Specifically, Figure 6 The structural block diagram of the data screening module 200 in the system provided by the embodiment of the application is shown.

[0102] In the preferred embodiment provided by the application, the data screening module 200 specifically comprises:

[0103] The data screening unit 201 is configured to analyze the historical interaction records, to screen out historical test question data having consistent attributes with the target test question data, and the consistent attributes include a question type category, a knowledge point label, and a test question use type;

[0104] The ratio calculation unit 202 is configured to obtain the number of modifications, the number of withdrawals and the total number of submissions of the target test question data and each piece of historical test question data, and calculate the ratio of the number of modifications to the total number of submissions and the ratio of the number of withdrawals to the total number of submissions respectively as a first ratio and a second ratio, and linearly weight the first ratio and the second ratio to obtain a corresponding withdrawal-modification ratio.

[0105] The curve construction unit 203 is configured to construct a withdrawal-modification ratio curve with time sequence as the horizontal axis and the withdrawal-modification ratio as the vertical axis.

[0106] Further, the copyright authentication accuracy optimization system based on the blockchain further comprises:

[0107] The first adjustment factor determination module 300 is configured to determine whether the withdrawal-modification ratio of the target test question data is located in the fitting interval of the withdrawal-modification ratio curve, and if yes, determine the first adjustment factor according to the change trend of the withdrawal-modification ratio curve.

[0108] Specifically, Figure 7 The structure block diagram of the first adjustment factor determination module 300 in the system provided by the embodiment of the present application is shown.

[0109] In the preferred embodiment provided by the present application, the first adjustment factor determination module 300 specifically comprises:

[0110] The ratio matching determination unit 301 is configured to substitute the withdrawal-modification ratio corresponding to the target test question data into the withdrawal-modification ratio curve, and determine whether the offset amplitude of the target test question data in the fitting interval is less than a preset threshold value.

[0111] The average slope calculation unit 302 is configured to calculate the average slope of the withdrawal-modification ratio curve if the offset amplitude is less than the preset threshold value, and take the average slope as the first adjustment factor.

[0112] Further, the copyright authentication accuracy optimization system based on the blockchain further comprises:

[0113] The second adjustment factor determination module 400 is configured to screen a plurality of representative test question data whose withdrawal-modification ratios are higher than a preset threshold value, extract semantic vectors of the modified fragments in the target test question data and each representative test question data based on semantic analysis, and calculate the semantic similarity in sequence, and determine the second adjustment factor based on the similarity result.

[0114] Specifically, Figure 8 The structure block diagram of the second adjustment factor determination module 400 in the system provided by the embodiment of the present application is shown.

[0115] In the preferred embodiment provided by the present application, the second adjustment factor determination module 400 specifically comprises:

[0116] Representative data extraction unit 401 is used to filter a preset number of representative test questions from historical test question data whose withdrawal and modification rate is higher than a preset threshold.

[0117] The semantic similarity calculation unit 402 is used to extract the semantic vectors of the target test question data and the modified segments in each representative test question data based on semantic analysis, and to calculate the semantic similarity between the target test question data and each representative test question data using cosine similarity.

[0118] The average value calculation unit 403 is used to calculate the average value of a preset number of semantic similarities and use the average value as a second adjustment factor.

[0119] Furthermore, the blockchain-based copyright authentication accuracy optimization system also includes:

[0120] The ownership credibility correction module 500 is used to correct the initial ownership credibility by combining the first adjustment factor and the second adjustment factor, and generate the corrected ownership credibility.

[0121] After determining the first adjustment factor and the second adjustment factor, based on the preset ownership credibility correction formula, the first adjustment factor and the second adjustment factor are combined with the initial ownership credibility to generate the corrected ownership credibility.

[0122] The ownership credibility correction formula is as follows: ,in This refers to the credibility of the revised ownership. This refers to the credibility of initial ownership. This refers to the first adjustment factor, which is the average slope of the withdrawal ratio curve. This refers to the adjustment coefficient corresponding to the first adjustment factor. This refers to the preset quantity. It refers to the first Semantic similarity of representative test questions This refers to the second adjustment factor, which is the average of a preset number of semantic similarities. This refers to the adjustment coefficient corresponding to the second adjustment factor.

[0123] It should be understood that, although the steps in the flowcharts of the embodiments of the present application are shown in a certain order according to the arrows, the steps are not necessarily executed in the order of the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in order, and the steps can be executed in other orders. Moreover, at least some of the steps in the embodiments can include a plurality of sub-steps or a plurality of stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the sub-steps or stages is not necessarily sequential, but can be round-robin or alternately executed with at least some of the other steps or sub-steps or stages of the other steps.

[0124] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a non-volatile computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of the methods. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0125] Any combination of the technical features of the above-mentioned embodiments can be made. In order to make the description concise, all possible combinations of the technical features in the above-mentioned embodiments are not described, but as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0126] The above embodiments only express several implementation manners of the present application, which are described in a more specific and detailed manner, but should not be understood as a limitation on the patent scope of the present application. It should be noted that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

[0127] The above only describes the preferred embodiments of the present application and should not be used to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A method for optimizing the accuracy of copyright authentication based on blockchain, characterized in that, The method includes: Acquire target test question data submitted and registered by users to the blockchain platform, obtain the initial ownership credibility determined for the target test question data, and obtain the user's historical interaction records in the blockchain platform; Analyze historical interaction records, extract historical test question data with the same attributes as the target test question data, calculate the withdrawal and modification ratio of the target test question data and each set of historical test question data, and construct a withdrawal ratio curve in chronological order; Determine whether the withdrawal and modification rate of the target test question data is within the fitted range of the withdrawal rate curve. If so, determine the first adjustment factor based on the changing trend of the withdrawal rate curve. Select a number of representative test questions with a withdrawal / modification rate higher than a preset threshold, extract semantic vectors of the modified segments in the target test questions and each representative test question based on semantic analysis, calculate semantic similarity in turn, and determine the second adjustment factor based on the similarity results; By combining the first adjustment factor and the second adjustment factor, the initial ownership credibility is corrected to generate the corrected ownership credibility.

2. The method for optimizing copyright authentication accuracy based on blockchain according to claim 1, characterized in that, The steps of analyzing historical interaction records, extracting historical test question data with attributes consistent with the target test question data, calculating the withdrawal / modification ratio of the target test question data and each set of historical test question data, and constructing a withdrawal ratio curve in chronological order include: Analyze historical interaction records and filter out historical test question data that have consistent attributes with the target test question data. Consistent attributes include question type, knowledge point tag, and test question purpose type. Obtain the number of modifications, withdrawals, and total submissions of the target test question data and each set of historical test question data. Calculate the ratio of the number of modifications to the total number of submissions and the ratio of the number of withdrawals to the total number of submissions, respectively, and use these as the first ratio and the second ratio. Perform a linear weighting on the first ratio and the second ratio to obtain the corresponding withdrawal and modification ratio. Construct a withdrawal rate curve with time sequence as the horizontal axis and withdrawal / modification rate as the vertical axis.

3. The method for optimizing copyright authentication accuracy based on blockchain according to claim 2, characterized in that, Determining whether the withdrawal / modification rate of the target test item data falls within the fitted range of the withdrawal rate curve, and if so, the steps for determining the first adjustment factor based on the trend of the withdrawal rate curve include: Substitute the withdrawal / modification ratio corresponding to the target test data into the withdrawal ratio curve to determine whether its offset within the fitting interval is less than the preset threshold. If the offset is less than the preset threshold, the average slope of the withdrawal ratio curve is calculated and used as the first adjustment factor.

4. The method for optimizing copyright authentication accuracy based on blockchain according to claim 3, characterized in that, The steps of selecting representative test questions with a withdrawal / modification rate exceeding a preset threshold, extracting semantic vectors from the target test questions and the modified segments in each representative test question based on semantic analysis, calculating semantic similarity sequentially, and determining the second adjustment factor based on the similarity results include: Select a predetermined number of representative test questions from historical test question data whose withdrawal / modification rate exceeds a predetermined threshold; Semantic vectors of the target test data and the modified segments in each representative test data are extracted based on semantic analysis, and cosine similarity is used to calculate the semantic similarity between the target test data and each representative test data. Calculate the average of a preset number of semantic similarities and use this average as a second adjustment factor.

5. The blockchain-based copyright authentication accuracy optimization method according to claim 4, characterized in that, After determining the first adjustment factor and the second adjustment factor, based on the preset ownership credibility correction formula, the first adjustment factor and the second adjustment factor are combined with the initial ownership credibility to generate the corrected ownership credibility. The ownership credibility correction formula is as follows: ,in This refers to the credibility of the revised ownership. This refers to the credibility of initial ownership. This refers to the first adjustment factor, which is the average slope of the withdrawal ratio curve. This refers to the adjustment coefficient corresponding to the first adjustment factor. This refers to the preset quantity. It refers to the first Semantic similarity of representative test questions This refers to the second adjustment factor, which is the average of a preset number of semantic similarities. This refers to the adjustment coefficient corresponding to the second adjustment factor.

6. A blockchain-based copyright authentication accuracy optimization system, characterized in that, The system includes: a data acquisition module, a data filtering module, a first adjustment factor determination module, a second adjustment factor determination module, and an ownership credibility correction module, wherein: The data acquisition module is used to acquire target test question data submitted and registered by users to the blockchain platform, and to acquire the initial ownership credibility determined for the target test question data, while also acquiring the user's historical interaction records in the blockchain platform; The data filtering module is used to parse historical interaction records, extract historical test question data with the same attributes as the target test question data, calculate the withdrawal and modification ratio of the target test question data and each set of historical test question data, and construct a withdrawal ratio curve in chronological order. The first adjustment factor determination module is used to determine whether the withdrawal and modification rate of the target test data is within the fitting range of the withdrawal rate curve. If so, the first adjustment factor is determined based on the changing trend of the withdrawal rate curve. The second adjustment factor determination module is used to screen a number of representative test question data with a withdrawal modification rate higher than a preset threshold, extract the semantic vectors of the modified segments in the target test question data and each representative test question data based on semantic analysis, calculate the semantic similarity in turn, and determine the second adjustment factor based on the similarity results. The ownership credibility correction module is used to correct the initial ownership credibility by combining the first adjustment factor and the second adjustment factor, and generate the corrected ownership credibility.

7. The blockchain-based copyright authentication accuracy optimization system according to claim 6, characterized in that, The data filtering module specifically includes: The data filtering unit is used to parse historical interaction records and filter out historical test question data that have the same attributes as the target test question data. The same attributes include question type, knowledge point tag and test question purpose type. The ratio calculation unit is used to obtain the number of modifications, withdrawals, and total submissions of the target test question data and each set of historical test question data, and to calculate the ratio of the number of modifications to the total number of submissions and the ratio of the number of withdrawals to the total number of submissions, respectively, as the first ratio and the second ratio. The first ratio and the second ratio are linearly weighted to obtain the corresponding withdrawal and modification ratio. The curve construction unit is used to construct a withdrawal rate curve with time sequence as the horizontal axis and withdrawal rate as the vertical axis.

8. The blockchain-based copyright authentication accuracy optimization system according to claim 7, characterized in that, The first adjustment factor determination module specifically includes: The ratio matching judgment unit is used to substitute the withdrawal and modification ratio corresponding to the target test data into the withdrawal ratio curve to determine whether its offset within the fitting interval is less than a preset threshold. The average slope calculation unit is used to calculate the average slope of the withdrawal ratio curve if the offset amplitude is less than a preset threshold, and to use the average slope as the first adjustment factor.

9. The blockchain-based copyright authentication accuracy optimization system according to claim 8, characterized in that, The second adjustment factor determination module specifically includes: The representative data extraction unit is used to filter a preset number of representative test questions from historical test question data whose withdrawal and modification rates exceed a preset threshold. The semantic similarity calculation unit is used to extract the semantic vectors of the target test question data and the modified segments in each representative test question data based on semantic analysis, and to calculate the semantic similarity between the target test question data and each representative test question data using cosine similarity. The average value calculation unit is used to calculate the average value of a preset number of semantic similarities and use this average value as a second adjustment factor.

10. The blockchain-based copyright authentication accuracy optimization system according to claim 9, characterized in that, After determining the first adjustment factor and the second adjustment factor, based on the preset ownership credibility correction formula, the first adjustment factor and the second adjustment factor are combined with the initial ownership credibility to generate the corrected ownership credibility. The ownership credibility correction formula is as follows: ,in This refers to the credibility of the revised ownership. This refers to the credibility of initial ownership. This refers to the first adjustment factor, which is the average slope of the withdrawal ratio curve. This refers to the adjustment coefficient corresponding to the first adjustment factor. This refers to the preset quantity. It refers to the first Semantic similarity of representative test questions This refers to the second adjustment factor, which is the average of a preset number of semantic similarities. This refers to the adjustment coefficient corresponding to the second adjustment factor.

Citation Information

Patent Citations

  • Intelligent questionnaire survey method and device, computer equipment and storage medium

    CN110070333A

  • Intelligent questionnaire data processing method based on cross check

    CN118505319A