Copyright authentication method, device, equipment, system and computer-readable storage medium
Through the online copyright certification method, digital works are reviewed and compared using processing strategies and review classification models, solving the problems of low efficiency and long cycle of copyright certification in the existing technology, and achieving fast and efficient copyright protection.
Patent Information
- Application Number
- CN201911093190.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-11-11
AI Technical Summary
In the prior art, copyright certification of works mainly relies on offline manual acceptance, resulting in time-consuming, high cost, low efficiency and long cycle, which cannot meet the demand for copyright certification in the Internet era.
Provide a copyright authentication method, by receiving copyright authentication requests for digital works, obtaining the digital works and types of works to be certified, determining the processing strategy and target review classification model, processing and reviewing the works to be certified, and if passed, it will be compared with the preset certified works library for copyright authentication.
It has realized online certification of copyright of digital works, and can complete copyright certification within the complexity of seconds, reduce labor costs, shorten copyright certification cycle, improve copyright certification efficiency, and promptly protect the copyright of the writers of the works.
Smart Images

Figure CN110781460B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial technology (Fintech), and in particular to a copyright authentication method, device, equipment, system and computer-readable storage medium. Background Art
[0002] With the development of computer technology, more and more technologies (big data, distributed, blockchain, artificial intelligence, etc.) are being applied in the financial field. The traditional financial industry is gradually transforming to financial technology (Fintech). However, due to the security and real-time requirements of the financial industry, higher requirements are also placed on technology.
[0003] At present, copyright certification of works is mainly completed through offline manual acceptance. The specific process is as follows: 1) The author of the work submits personal information and works to the intellectual property agent; 2) The intellectual property agent determines whether the work meets the registration conditions and determines the type of copyright registration; 3) If the work meets the copyright application conditions, the intellectual property agent submits the work registration application form to the Copyright Center; 4) After receiving the application, the Copyright Center reviews the application materials and decides whether to issue the copyright registration certificate for the work. The entire process takes about 20 to 30 working days, which is a long cycle. With the popularization of Internet technology, hundreds of thousands of original digital works are generated on the Internet every day. The offline manual acceptance of copyright certification of works has been unable to meet the current needs due to its time-consuming, labor-intensive, high cost, low efficiency and long cycle, resulting in a large number of infringing and pirated works being spread on the Internet. Therefore, there is an urgent need for a copyright certification method for digital works to shorten the copyright certification cycle, improve the efficiency of copyright certification, and protect the copyright of the author of the work in a timely manner. Summary of the invention
[0004] The main purpose of the present invention is to provide a copyright authentication method, device, equipment, system and computer-readable storage medium, aiming to shorten the copyright authentication cycle and improve the efficiency of copyright authentication.
[0005] To achieve the above object, the present invention provides a copyright authentication method, which comprises:
[0006] Upon receiving a digital work copyright authentication request, obtaining the digital work to be authenticated and the type of the work according to the digital work copyright authentication request;
[0007] Determine a corresponding processing strategy and a target review classification model according to the type of the work, and process the digital work to be authenticated based on the processing strategy to obtain a target input object;
[0008] Input the target input object into the target review classification model to obtain a review result, and determine whether the review is passed based on the review result;
[0009] When the review is passed, the digital work to be authenticated is compared with the authenticated digital works in the preset authentication work library for copyright authentication.
[0010] Optionally, the step of determining a corresponding processing strategy and a target review classification model according to the work type, and processing the digital work to be authenticated based on the processing strategy to obtain a target input object includes:
[0011] If the type of the work is a text work, the corresponding processing strategy is determined to be the first processing strategy, and the target review classification model is determined to be the first review classification model;
[0012] Performing word segmentation processing on the digital work to be authenticated based on the first processing strategy to obtain a first word segmentation text;
[0013] Inputting the first segmented word text into a preset word vector model to obtain a first word vector for each segmented word in the first segmented word text;
[0014] A first document vector corresponding to the digital work to be authenticated is obtained according to the first word vector, wherein the target input object is the first document vector.
[0015] Optionally, the step of determining a corresponding processing strategy and a target review classification model according to the work type, and processing the digital work to be authenticated based on the processing strategy to obtain a target input object includes:
[0016] If the work type is a picture work, determining the corresponding processing strategy to be the second processing strategy, and determining the target review classification model to be the second review classification model;
[0017] The digital work to be authenticated is preprocessed based on the second processing strategy to obtain an input image, wherein the preprocessing includes scaling processing and grayscale processing, and the target input object is the input image.
[0018] Optionally, the step of determining a corresponding processing strategy and a target review classification model according to the work type, and processing the digital work to be authenticated based on the processing strategy to obtain a target input object includes:
[0019] If the work type is an audio work, determining the corresponding processing strategy to be the third processing strategy, and determining the target review classification model to be the third review classification model;
[0020] Converting the digital work to be authenticated into a text work type based on the third processing strategy to obtain a converted digital work to be authenticated;
[0021] Performing word segmentation processing on the converted digital work to be authenticated to obtain a second word segmentation text;
[0022] Inputting the second segmented word text into a preset word vector model to obtain a second word vector for each segmented word in the second segmented word text;
[0023] A second document vector corresponding to the converted digital work to be authenticated is obtained according to the second word vector, wherein the target input object is the second document vector.
[0024] Optionally, the target review classification model includes multiple ones, the number of the review results is the same as the number of the target review classification models, and the step of judging whether the review is passed based on the review results includes:
[0025] Check whether the multiple review results are all qualified;
[0026] If all of the above review results are qualified, the review is considered to be passed;
[0027] If at least one of the multiple review results is unqualified, the review is determined to be unsuccessful.
[0028] Optionally, if the work type is a text work, the step of comparing the digital work to be authenticated with authenticated digital works in a preset authentication work library to perform copyright authentication includes:
[0029] Calculating a first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library by using a preset document search engine;
[0030] Filtering a first preset number of similar text works from the authenticated text works according to the first similarity value;
[0031] Calculating a first longest common subsequence between the similar text work and the digital work to be authenticated, and calculating a length ratio between the similar text work and the digital work to be authenticated based on the length of the first longest common subsequence to obtain a first calculation result;
[0032] Detecting whether there is a length ratio greater than a first preset threshold in the first calculation result;
[0033] If there is a length ratio greater than the first preset threshold, it is determined that the copyright authentication fails;
[0034] If there is no length ratio greater than the first preset threshold, it is determined that the copyright authentication is passed.
[0035] Optionally, the step of calculating the first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library by using a preset document search engine includes:
[0036] Performing word segmentation processing on the digital work to be authenticated by using a preset document search engine to obtain a word segmentation set;
[0037] Performing an inverted index on the certified text works in the preset certified works library through the preset document search engine, and calculating the score corresponding to each word in the word set according to the inverted index result;
[0038] The scores of the segmented words are summed up to obtain a first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library.
[0039] Optionally, if the work type is a picture work, the step of comparing the digital work to be authenticated with authenticated digital works in a preset authentication work library to perform copyright authentication includes:
[0040] Calculating a second similarity value between the digital work to be authenticated and authenticated image works in a preset authentication work library by using a preset image retrieval engine;
[0041] Filtering a second preset number of similar picture works from the authenticated picture works according to the second similarity value;
[0042] Extracting a first scale-invariant feature transform (SIFT) feature vector of the digital work to be authenticated, and extracting a second SIFT feature vector of the similar image work;
[0043] Calculating a cosine distance between the first SIFT feature vector and the second SIFT feature vector to obtain a second calculation result;
[0044] Detecting whether there is a cosine distance greater than a second preset threshold in the second calculation result;
[0045] If there is a cosine distance greater than the second preset threshold, it is determined that the copyright authentication fails;
[0046] If there is no cosine distance greater than the second preset threshold, it is determined that the copyright authentication is passed.
[0047] Optionally, if the work type is an audio work, the step of comparing the digital work to be authenticated with authenticated digital works in a preset authentication work library to perform copyright authentication includes:
[0048] Converting the digital work to be authenticated into a text work type to obtain an audio text work to be authenticated;
[0049] Calculating a third similarity value between the audio text work to be authenticated and authenticated audio text works in a preset authentication work library through a preset document search engine;
[0050] Retrieving a third preset number of similar audio text works from the authenticated audio text works according to the third similarity value;
[0051] Calculating the second longest common subsequence between the similar audio file work and the audio text work to be authenticated, and calculating the length ratio between the similar audio text work and the audio text work to be authenticated according to the length of the second longest common subsequence, to obtain a third calculation result;
[0052] Detecting whether there is a length ratio greater than a third preset threshold in the third calculation result;
[0053] If there is a length ratio greater than the third preset threshold, it is determined that the copyright authentication fails;
[0054] If there is no length ratio greater than the third preset threshold, it is determined that the copyright authentication is passed.
[0055] Optionally, the copyright authentication method further includes:
[0056] When the copyright authentication is passed, obtaining the work information of the digital work to be authenticated;
[0057] Generate corresponding copyright authentication information based on the work information, and generate a data upload request based on the copyright authentication information;
[0058] The data chain upload request is sent to the copyright authentication alliance chain, so that the copyright authentication alliance chain completes the chain upload operation of the digital work to be authenticated based on the data chain upload request.
[0059] In addition, to achieve the above-mentioned purpose, the present invention also provides a copyright authentication device, the copyright authentication device comprising:
[0060] A first acquisition module is used to, upon receiving a digital work copyright authentication request, acquire the digital work to be authenticated and the work type according to the digital work copyright authentication request;
[0061] A processing module, used to determine a corresponding processing strategy and a target review classification model according to the type of the work, and process the digital work to be authenticated based on the processing strategy to obtain a target input object;
[0062] A review module, used for inputting the target input object into the target review classification model, obtaining a review result, and judging whether the review is passed based on the review result;
[0063] The copyright authentication module is used to compare the digital work to be authenticated with the authenticated digital works in the preset authentication work library to perform copyright authentication when the review is passed.
[0064] In addition, to achieve the above-mentioned purpose, the present invention also provides a copyright authentication device, which includes: a memory, a processor, and a copyright authentication program stored in the memory and executable on the processor, and the copyright authentication program implements the steps of the copyright authentication method described above when executed by the processor.
[0065] In addition, to achieve the above-mentioned purpose, the present invention also provides a copyright authentication system, which includes a copyright authentication device and a copyright authentication alliance chain; wherein:
[0066] The copyright authentication device is the copyright authentication device as described above;
[0067] The copyright authentication alliance chain is used to receive a data upload request sent by the copyright authentication device;
[0068] Based on the data on-chain request, the data information to be on-chain is obtained, and based on the consensus algorithm, the on-chain operation of the data information to be on-chain is completed.
[0069] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a copyright authentication program is stored, and when the copyright authentication program is executed by a processor, the steps of the copyright authentication method described above are implemented.
[0070] The present invention provides a copyright authentication method, device, equipment, system and computer-readable storage medium. When receiving a digital work copyright authentication request, the digital work to be authenticated and the work type are obtained according to the digital work copyright authentication request; the corresponding processing strategy and target review classification model are determined according to the work type, and the digital work to be authenticated is processed based on the processing strategy to obtain the target input object; the target input object is input into the target review classification model to obtain the review result, and whether the review is passed is determined based on the review result; and when the review is passed, the digital work to be authenticated is compared with the authenticated digital works in the preset authentication work library to perform copyright authentication. Through the above method, online authentication of the copyright of digital works can be realized, and copyright authentication can be completed within a time complexity of seconds for different types of digital works. Compared with the copyright authentication performed manually in the prior art, the present invention can reduce labor costs, shorten the copyright authentication cycle, and improve the efficiency of copyright authentication, so that the copyright of the author of the work can be protected in time. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1A schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention;
[0072] Figure 2 It is a flowchart of the first embodiment of the copyright authentication method of the present invention;
[0073] Figure 3 A schematic diagram of the system structure of the copyright authentication system of the present invention;
[0074] Figure 4 Schematic diagram of the functional modules of the first embodiment of the copyright authentication device of the present invention.
[0075] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0076] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0077] Reference Figure 1 , Figure 1 The figure is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention.
[0078] The copyright authentication device in the embodiment of the present invention may be a smart phone, or a terminal device such as a PC (Personal Computer), a tablet computer, or a portable computer.
[0079] like Figure 1 As shown, the copyright authentication device may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory, or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.
[0080] Those skilled in the art will understand that Figure 1 The structure of the copyright authentication device shown in the figure does not constitute a limitation on the copyright authentication device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0081] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a copyright authentication program.
[0082] exist Figure 1 In the terminal shown, the network interface 1004 is mainly used to connect to the background server and communicate data with the background server; the user interface 1003 is mainly used to connect to the client and communicate data with the client; and the processor 1001 can be used to call the copyright authentication program stored in the memory 1005 and execute the various steps of the following copyright authentication method.
[0083] Based on the above hardware structure, various embodiments of the copyright authentication method of the present invention are proposed.
[0084] The invention provides a copyright authentication method.
[0085] Reference Figure 2 , Figure 2 It is a flowchart of the first embodiment of the copyright authentication method of the present invention.
[0086] In this embodiment, the copyright authentication method includes:
[0087] Step S10, when receiving a digital work copyright authentication request, obtaining the digital work to be authenticated and the work type according to the digital work copyright authentication request;
[0088] In this embodiment, the copyright authentication method is applied to a copyright authentication system, which includes a copyright authentication device and a copyright authentication alliance chain, wherein the copyright authentication method of this embodiment is implemented by a copyright authentication device, which is equipped with a copyright authentication system. The copyright authentication alliance chain can be composed of a copyright authentication agency node, a notary agency node, a judicial agency node, and an external node, which is used to receive a data chain request sent by the copyright authentication device, and then obtain the data information to be chained based on the data chain request, and then complete the chain operation of the data information to be chained based on the consensus algorithm, that is, to realize copyright authentication based on blockchain.
[0089] In this embodiment, when a user needs to authenticate the copyright of his / her work, he / she can upload his / her digital work (such as text work, picture work and audio work, etc.) through the corresponding software of the user terminal (such as a PC, a smart phone, etc.), and fill in relevant information (including but not limited to the type of work, information of the author, etc.), thereby triggering a request for copyright authentication of the digital work. At this time, when the copyright authentication system receives the request for copyright authentication of the digital work, it obtains the digital work to be authenticated and the type of work according to the request for copyright authentication of the digital work. Of course, it can be understood that in specific implementation, when triggering the request for copyright authentication of the digital work, the user can only upload his / her digital work, and after obtaining the digital work to be authenticated, the copyright authentication system can determine the corresponding type of work according to its format.
[0090] Step S20, determining a corresponding processing strategy and a target review classification model according to the type of the work, and processing the digital work to be authenticated based on the processing strategy to obtain a target input object;
[0091] After obtaining the digital work to be authenticated and the type of work, the corresponding processing strategy and target review classification model are determined according to the type of work, and the digital work to be authenticated is processed based on the processing strategy to obtain the target input object.
[0092] Specifically, if the work type is a text work, the corresponding processing strategy is determined to be the first processing strategy, and the target review classification model is determined to be the first review classification model; then, based on the first processing strategy, the digital work to be authenticated is segmented to obtain a first segmented text; the first segmented text is input into a preset word vector model to obtain a first word vector for each segmented word in the first segmented text; and then, based on the first word vector, a first document vector corresponding to the digital work to be authenticated is obtained, wherein the target input object is the first document vector.
[0093] If the work type is a picture work, the corresponding processing strategy is determined to be the second processing strategy, and the target review classification model is determined to be the second review classification model; based on the second processing strategy, the digital work to be authenticated is preprocessed to obtain an input picture, wherein the preprocessing includes scaling processing and grayscale processing, and the target input object is the input picture.
[0094] If the work type is an audio work, the corresponding processing strategy is determined to be the third processing strategy, and the target review classification model is determined to be the third review classification model; based on the third processing strategy, the digital work to be authenticated is converted into a text work type to obtain the converted digital work to be authenticated; then it is processed according to the processing method of the text work, that is, the converted digital work to be authenticated is segmented to obtain a second segmented text; then the second segmented text is input into a preset word vector model to obtain a second word vector for each segmented word in the second segmented text; according to the second word vector, a second document vector corresponding to the converted digital work to be authenticated is obtained, wherein the target input object is the second document vector.
[0095] The specific execution process may refer to the second embodiment described below and will not be described in detail here.
[0096] Step S30, inputting the target input object into the target review classification model to obtain a review result, and judging whether the review is passed based on the review result;
[0097] After obtaining the target input object, the target input object is input into the target review classification model to obtain the review result, and whether the review is passed is determined based on the review result. Among them, during the review, the main purpose is to review whether the works submitted by the user involve bad information, and the corresponding target review classification model can include 6 categories. When the target review classification model can include multiple, the number of review results is the same as the number of target review classification models, that is, the corresponding review results also include multiple. When judging whether the review is passed based on multiple review results, it is necessary to detect whether the multiple review results are all qualified; if the multiple review results are all qualified, it means that the digital work to be certified does not involve bad information, and the review is determined to be passed; if at least one of the multiple review results is unqualified, it means that the digital work to be certified involves bad information, and the review is determined to be failed.
[0098] Of course, in a specific embodiment, a review classification can be constructed for each type of digital work. Correspondingly, there is only one review result. In this case, it is only necessary to check whether the review result is qualified.
[0099] Step S40, when the review is passed, the digital work to be authenticated is compared with authenticated digital works in a preset authentication work library to perform copyright authentication.
[0100] When the review is passed, the digital work to be authenticated is compared with the authenticated digital works in the preset authentication work library for copyright authentication. Specifically, different copyright authentication methods need to be used for different types of digital works to be authenticated. The specific copyright authentication process can refer to the fourth embodiment described below, which will not be described in detail here. By comparing the digital work to be authenticated with the authenticated digital works in the preset authentication work library, it is possible to detect whether there is a reference / plagiarism relationship between the work to be authenticated and the authenticated work, so as to determine whether to authenticate the version.
[0101] The embodiment of the present invention provides a copyright authentication method. When a digital work copyright authentication request is received, the digital work to be authenticated and the work type are obtained according to the digital work copyright authentication request; the corresponding processing strategy and target review classification model are determined according to the work type, and the digital work to be authenticated is processed based on the processing strategy to obtain the target input object; the target input object is input into the target review classification model to obtain the review result, and whether the review is passed is determined based on the review result; and when the review is passed, the digital work to be authenticated is compared with the authenticated digital works in the preset authentication work library to perform copyright authentication. Through the above method, online authentication of the copyright of digital works can be realized, and copyright authentication can be completed within a time complexity of seconds for different types of digital works. Compared with the prior art of manual copyright authentication, the embodiment of the present invention can reduce labor costs, shorten the copyright authentication cycle, and improve the efficiency of copyright authentication, so that the copyright of the author of the work can be protected in time.
[0102] Further, based on Figure 2 The first embodiment shown here proposes a second embodiment of the copyright authentication method of the present invention.
[0103] In this embodiment, as one implementation method, step S20 may include:
[0104] Step a11, if the work type is a text work, determining the corresponding processing strategy to be the first processing strategy, and determining the target review classification model to be the first review classification model;
[0105] Step a12, performing word segmentation processing on the digital work to be authenticated based on the first processing strategy to obtain a first word segmentation text;
[0106] Step a13, inputting the first segmented word text into a preset word vector model to obtain a first word vector of each segmented word in the first segmented word text;
[0107] Step a14, obtaining a first document vector corresponding to the digital work to be authenticated according to the first word vector, wherein the target input object is the first document vector.
[0108] In this embodiment, the processing process for the text work is as follows:
[0109] If the type of work is a text work, the corresponding processing strategy is determined to be the first processing strategy, and the target review classification model is determined to be the first review classification model. Among them, the first review classification model is pre-trained, and its type can be a binary classification model such as an SVM (Support Vector Machine) model, a Bayesian model, a logistic regression model, a convolutional neural network model, etc. The following training process is explained with the SVM model. The first review classification model also includes 6 categories, and its training process is: take 50,000 annotated text works involving and 50,000 not involving bad information, respectively, and after word segmentation for each text work, obtain the word vector through a preset word vector model (optionally a word2vec model), and then add the word vectors according to the corresponding dimensions to obtain the document vector corresponding to each text work, and then train an SVM classification model based on the above 100,000 document vectors.
[0110] After the first processing strategy is determined, the digital work to be authenticated is segmented based on the first processing strategy to obtain a first segmented text, wherein the segmentation process may use a preset tool, such as NLPIR of the Chinese Academy of Sciences, LTP of Harbin Institute of Technology, Jieba segmentation, etc. The specific segmentation process is consistent with the prior art and will not be described in detail here.
[0111] Then, the first segmented text is input into the preset word vector model to obtain the first word vector of each segmented word in the first segmented text. Among them, the preset word vector model can be word2vec (word to vector, a related model used to generate word vectors). Word2vec maps each Chinese word to a high-dimensional vector (usually a 200-dimensional vector), and for any two Chinese words, the closer they are semantically, the closer the vector distance obtained after mapping. Therefore, the semantic similarity of Chinese words can be described based on the distance between word vectors.
[0112] Finally, the first document vector corresponding to the digital work to be authenticated is obtained based on the first word vector, wherein the target input object is the first document vector, that is, the first document vector is subsequently input into the corresponding first review classification model to obtain the review result. The first document vector is obtained by adding the first word vector according to the corresponding dimension to obtain the corresponding first document vector.
[0113] Through the above method, digital works of text works can be processed to obtain corresponding target input objects, so as to facilitate subsequent input into the target review classification model to obtain review results.
[0114] As another implementation, step S20 may further include:
[0115] Step a21, if the work type is a picture work, determining the corresponding processing strategy to be the second processing strategy, and determining the target review classification model to be the second review classification model;
[0116] Step a22, pre-processing the digital work to be authenticated based on the second processing strategy to obtain an input image, wherein the pre-processing includes scaling processing and grayscale processing, and the target input object is the input image.
[0117] In this embodiment, the processing process for the picture work is as follows:
[0118] If the type of work is a picture work, the corresponding processing strategy is determined to be the second processing strategy, and the target review classification model is determined to be the second review classification model. Among them, the second review classification model is pre-trained, and its type can be optionally a classification model based on a convolutional neural network. The second review classification model also includes 6 categories, and its training process is: take 50,000 labeled pictures involving and 50,000 pictures not involving bad information, and pre-process each picture work. The pre-processing process includes scaling and grayscale processing. Among them, scaling is to scale the size of the picture to a preset size, such as 128 pixels * 128 pixels, and grayscale processing is to convert the scaled picture into a grayscale picture, and then train a classification model based on a convolutional neural network based on the above 100,000 pre-processed pictures.
[0119] After determining that the second processing strategy is obtained, the digital work to be authenticated is preprocessed based on the second processing strategy to obtain an input image, wherein the preprocessing includes scaling processing and grayscale processing. The scaling processing is to scale the size of the image to a preset size, such as 128 pixels * 128 pixels, and the grayscale processing is to convert the scaled image into a grayscale image. The target input object is the input image, that is, the input image is subsequently input into the corresponding second review classification model to obtain a review result.
[0120] Through the above method, digital works such as picture works can be processed to obtain corresponding target input objects, so as to facilitate subsequent input into the target review classification model to obtain review results.
[0121] As another implementation, step S20 may further include:
[0122] Step a31, if the work type is an audio work, determining the corresponding processing strategy to be the third processing strategy, and determining the target review classification model to be the third review classification model;
[0123] Step a32, converting the digital work to be authenticated into a text work type based on the third processing strategy to obtain a converted digital work to be authenticated;
[0124] Step a33, performing word segmentation processing on the converted digital work to be authenticated to obtain a second word segmentation text;
[0125] Step a34, inputting the second segmented word text into a preset word vector model to obtain a second word vector for each segmented word in the second segmented word text;
[0126] Step a35, obtaining a second document vector corresponding to the converted digital work to be authenticated according to the second word vector, wherein the target input object is the second document vector.
[0127] In this embodiment, the processing process for the audio work is as follows:
[0128] If the work type is an audio work, the corresponding processing strategy is determined to be the third processing strategy, and the target review classification model is determined to be the third review classification model. The third review classification model is pre-trained, and its type can be a binary classification model such as a SVM (Support Vector Machine) model, a Bayesian model, a logistic regression model, a convolutional neural network model, etc. The third review classification model can be the same as the first review classification model, or it can be another type of binary classification model trained based on the training method of the first review classification model.
[0129] After determining that the third processing strategy is obtained, the digital work to be authenticated is first converted into a text work type based on the third processing strategy to obtain the converted digital work to be authenticated. Specifically, the audio work can be converted into a text work type through a speech recognition tool. Then, the converted digital work to be authenticated is segmented to obtain a second segmented text, wherein the segmentation process can use preset tools, such as NLPIR of the Chinese Academy of Sciences, LTP of Harbin Institute of Technology, and Jieba segmentation. The specific segmentation process is consistent with the prior art and will not be repeated here.
[0130] Then, the second segmented word text is input into a preset word vector model to obtain a second word vector of each segmented word in the second segmented word text. The preset word vector model may optionally be word2vec (word to vector, a related model for generating word vectors).
[0131] Finally, the second document vector corresponding to the digital work to be authenticated is obtained based on the second word vector, wherein the target input object is the second document vector, that is, the second document vector is subsequently input into the third review classification model corresponding to the value to obtain the review result. The second document vector is obtained by adding the second word vector according to the corresponding dimension to obtain the corresponding second document vector.
[0132] Through the above method, digital works of audio works can be processed to obtain corresponding target input objects, so as to facilitate subsequent input into the target review classification model to obtain review results.
[0133] Based on the above first and second embodiments, a third embodiment of the copyright authentication method of the present invention is proposed.
[0134] In this embodiment, the target review classification model includes multiple ones, the number of the review results is the same as the number of the target review classification models, and the step of "determining whether the review is passed based on the review results" includes:
[0135] Step b1, detecting whether the plurality of review results are all qualified;
[0136] Step b2: if the plurality of review results are all qualified, the review is determined to be passed;
[0137] Step b3: If at least one of the multiple review results is unqualified, the review is determined to be unsuccessful.
[0138] In this embodiment, the target review classification model may include multiple, the number of review results is the same as the number of target review classification models, and the corresponding review results also include multiple. The judgment process for whether the review is passed is: detecting whether multiple review results are all qualified.
[0139] If all of the above review results are qualified, it means that the digital work to be authenticated does not involve any harmful information, and the review is considered to have passed;
[0140] If at least one of the multiple review results is unqualified, it means that the digital work to be authenticated contains bad information, and the review is determined to be unqualified.
[0141] Furthermore, based on the above-mentioned first embodiment, a fourth embodiment of the copyright authentication method of the present invention is proposed.
[0142] In this embodiment, if the work type is a text work, step S40 includes:
[0143] Step c11, calculating a first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library by using a preset document search engine;
[0144] Step c12, screening a first preset number of similar text works from the authenticated text works according to the first similarity value;
[0145] Step c13, calculating the first longest common subsequence between the similar text work and the digital work to be authenticated, and calculating the length ratio between the similar text work and the digital work to be authenticated according to the length of the first longest common subsequence, to obtain a first calculation result;
[0146] Step c14, detecting whether there is a length ratio greater than a first preset threshold in the first calculation result;
[0147] Step c15: if there is a length ratio greater than the first preset threshold, it is determined that the copyright authentication fails;
[0148] Step c16: If there is no length ratio greater than the first preset threshold, it is determined that the copyright authentication is passed.
[0149] Wherein, step c11 includes:
[0150] Step c111, performing word segmentation processing on the digital work to be authenticated by using a preset document search engine to obtain a word segmentation set;
[0151] Step c112, performing an inverted index on the certified text works in the preset certified works library through the preset document search engine, and calculating the score corresponding to each word in the word set according to the inverted index result;
[0152] Step c113, summing up the scores of each word segmentation to obtain a first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library.
[0153] In this embodiment, if the work type is a text work, the corresponding copyright authentication process is as follows:
[0154] The first similarity value between the digital work to be authenticated and the authenticated text works in the preset authentication works library is calculated through a preset document search engine. Among them, the preset document search engine can optionally be an ES (Elastic Search) search engine. ES is a distributed, highly scalable, and highly real-time search and data analysis engine, which can easily enable large amounts of data to have the ability to search, analyze, and explore. Specifically, the digital work to be authenticated is first segmented through the ES search engine to obtain a segmentation set, wherein the ES search engine has its own segmenter, which can perform segmentation on the digital work to be authenticated. Then, the authenticated text works in the preset authentication works library are inverted indexed through the ES search engine, and the scores corresponding to each segmentation in the segmentation set are calculated based on the inverted index results. When the ES search engine is used to perform an inverted index on the certified text works, the ES search engine's built-in word segmenter is also used to perform word segmentation first, and the word frequency information and position information of each word segmentation will be obtained, and then an inverted index between each word segmentation and the work document will be established, where the inverted index is a dictionary-like data structure (key-value), the key of the dictionary is a word segmentation, and the value is a list of works containing the word segmentation, as well as the position information and word frequency information of the word segmentation in each work. Through the inverted index, the document list and word frequency information containing the word segmentation can be quickly obtained based on the word segmentation. When calculating the score corresponding to each word segmentation in the word segmentation set based on the inverted index results, the document containing the word segmentation, the word frequency and inverse document frequency of the word segmentation in the document can be obtained based on the inverted index results, and then the score of each word segmentation can be calculated based on the word frequency and inverse document frequency. For example, after the digital works to be certified are segmented, a segmentation set including the segmentations "A", "B", "C", "D", "E" and "F" is obtained. First, the work set including the segmentation "A" is found based on the inverted index of the certified text works, and the score corresponding to the segmentation "A" is calculated. Specifically, the score of each work in the work set can be calculated first. The score of each work can be the product of the word frequency of the segmentation A in the work and the inverse document frequency of the word A (of course, other calculation methods can also be set according to actual conditions, such as calculation based on word frequency, inverse document frequency and position information), and then the score of each work is added to obtain the score corresponding to the segmentation "A"; after obtaining the score corresponding to the segmentation "A", the same operation is performed on the segmentations "B", "C", "D", "E" and "F" to obtain the scores corresponding to the segmentations "B", "C", "D", "E" and "F" respectively. In addition, it can be understood that the certified text works in the preset certified works library can be stored in advance according to a specific index structure using the ES search engine. In this case, there is no need to execute the step of "performing inverse indexing of the certified text works in the preset certified works library through the preset document search engine".After obtaining the scores corresponding to the participles, the scores of the participles are summed up to obtain a first similarity value between the digital work to be authenticated and the authenticated text works in the preset authentication work library.
[0155] After obtaining the first similarity value between the digital work to be authenticated and the authenticated text work, a first preset number of similar text works are screened from the authenticated text works according to the first similarity value, wherein the first preset number can be set according to actual needs, for example, it can be set to a preset value, such as 1000, or it can be set to the number of first similarity values greater than a preset value, which is not specifically limited here. During screening, the first similarity values are sorted in order from large to small, and the first preset number of text works ranked in the front are taken as similar text works. It should be noted that the premise of considering the existence of plagiarism relationship between two text works is that the words used in the two works are consistent, and the words before and after each word are also consistent. If the words used in the two works are not the same, then there will definitely be no plagiarism relationship between the two news. Therefore, here, the ES search engine is used to preliminarily determine whether the words are consistent, so as to perform preliminary screening, the purpose of which is to screen out text works containing all or most of the words in the text work to be authenticated from the authenticated text works, so as to narrow the scope of authentication comparison, save server resources, and further improve the efficiency of copyright authentication of text works. In addition, it should be noted that, in specific implementation, of course, the similarity value between the digital work to be authenticated and the authenticated text works in the preset authentication work library can be directly calculated based on the calculation of the longest announcement subsequence, and then whether the copyright authentication is passed can be determined based on the similarity value. However, in comparison, the complexity of the calculation process of the initial screening through the ES search engine is obviously lower than the calculation complexity of the longest common subsequence. With the help of the ES search engine, it takes milliseconds to retrieve the top 1,000 works with the same similarity value as the digital work to be authenticated from tens of millions of authenticated text works. Therefore, in this embodiment, the ES search engine is first used for initial screening, and then combined with the subsequent calculation of the longest common subsequence. Compared with the method of directly calculating the similarity value based on the longest common subsequence, the efficiency of copyright authentication of works can be further improved.
[0156] Then, the first longest common subsequence between the similar text work and the digital work to be authenticated is calculated, where the "first" in the first longest common subsequence has no real meaning and is only used to distinguish it from the subsequent second longest announcement subsequence. The longest common subsequence is to find the longest common subsequence of two sequences. For example, given two strings X =<x1,x2,x3,…,xn> , Y=<y1,y2,y3,…,ym> , there are two subscript sequences of length k<i1,i2,…,ik> ,<j1,j2,…,jk> , so that the characters of the strings X and Y at the positions corresponding to the subscripts i and j are equal, and the substring corresponding to the longest subscript sequence that meets the above requirements is the longest common subsequence between X and Y. For example, the longest common subsequence between the string ABCDE and the string XAYCDZ is ACD. Then, the length ratio between the similar text work and the digital work to be authenticated is calculated according to the length of the first longest common subsequence to obtain the first calculation result. The number of the first longest common subsequences corresponds to the same number of similar text works. For the calculation of the length ratio, for example, when the first longest common subsequence of a similar text work includes 100 characters, its length is 100. If the digital work to be authenticated includes 1000 characters, the length ratio is 100 / 1000=0.1.
[0157] Then, it is detected whether there is a length ratio greater than a first preset threshold value in the first calculation result, wherein the first preset threshold value can be set according to actual needs, for example, it can be set to 0.8, and is not specifically limited here. If there is a length ratio greater than the first preset threshold value, it is considered that there is a reference / plagiarism relationship, and the copyright authentication is determined to be unsuccessful; if there is no length ratio greater than the first preset threshold value, the copyright authentication is determined to be successful.
[0158] It should be noted that, in specific implementation, the processes of "calculating the first longest common subsequence between similar textual works and the digital work to be authenticated, and calculating the length ratio between the similar textual works and the digital work to be authenticated based on the length of the first longest common subsequence" and "detecting whether there is a length ratio greater than a first preset threshold" can be performed simultaneously, that is, the first longest common subsequence between each similar textual work and the digital work to be authenticated is calculated in turn, and then the length ratio between the similar textual works and the digital work to be authenticated is calculated based on the length of the first longest common subsequence, and then whether the length ratio is greater than the first preset threshold. Once it is detected that the length ratio corresponding to a certain similar textual work is greater than the first preset threshold, it is considered that there is a citation / plagiarism relationship, and the copyright authentication can be determined to be unsuccessful. At this time, there is no need to calculate the first longest common subsequence between other similar textual works and the digital work to be authenticated and subsequent steps, which can save server resources and further improve the efficiency of copyright authentication.
[0159] Through the above method, it is possible to detect whether there is a citation / plagiarism relationship between the work to be authenticated and the authenticated work, so as to determine whether to authenticate the version.
[0160] In this embodiment, if the work type is a picture work, step S40 may further include:
[0161] Step c21, calculating a second similarity value between the digital work to be authenticated and authenticated image works in a preset authentication work library by using a preset image retrieval engine;
[0162] Step c22, screening out a second preset number of similar picture works from the authenticated picture works according to the second similarity value;
[0163] Step c23, extracting a first scale-invariant feature transform (SIFT) feature vector of the digital work to be authenticated, and extracting a second SIFT feature vector of the similar image work;
[0164] Step c24, calculating the cosine distance between the first SIFT feature vector and the second SIFT feature vector to obtain a second calculation result;
[0165] Step c25, detecting whether there is a cosine distance greater than a second preset threshold in the second calculation result;
[0166] Step c26, if there is a cosine distance greater than the second preset threshold, it is determined that the copyright authentication fails;
[0167] Step c27: If there is no cosine distance greater than the second preset threshold, it is determined that the copyright authentication is passed.
[0168] In this embodiment, if the work type is a picture work, the corresponding copyright authentication process is as follows:
[0169] The second similarity value between the digital work to be authenticated and the authenticated picture works in the preset authentication work library is calculated by the preset image retrieval engine. Among them, the preset image retrieval engine can be optionally a CBIR (Content-based image retrieval) engine, and the core of the CBIR engine is to use the visual features of the image to retrieve the image. In essence, it is an approximate matching technology that integrates the technical achievements of multiple fields such as computer vision, image processing, image understanding and database, in which the feature extraction and index establishment can be automatically completed by the computer, avoiding the subjectivity of manual description. The user retrieval process generally provides a sample image (Query by Example) or draws a sketch (Query by Sketch), the system extracts the features of the query image, and then compares it with the features in the database, and returns the image similar to the query feature to the user. It should be noted that since the authenticated picture works in the preset authentication work library are stored in preset sizes, it is necessary to perform corresponding scaling processing on the digital works to be authenticated of the picture type to obtain the digital works to be authenticated of the same preset size, and then calculate the second similarity value between the scaled digital works to be authenticated and the authenticated picture works in the preset authentication work library through the preset image retrieval engine.
[0170] After obtaining the second similarity value between the digital work to be authenticated and the authenticated picture work, a second preset number of similar picture works are screened from the authenticated picture works according to the second similarity value, wherein the second preset number may be the same as or different from the first preset number, and may be set according to actual needs, and is not specifically limited here. During screening, the second similarity values are sorted in order from large to small, and the second preset number of text works ranked in the front are taken as similar text works. It should be noted that the premise of considering the existence of a plagiarism relationship between two picture works is that the low-level features such as color, shape, texture, etc. of the two picture works are consistent, and the arrangement of each feature is also consistent. If the low-level features of the two picture works are not similar, then there will definitely be no plagiarism relationship between the two picture works. Therefore, the purpose of preliminary screening by the CBIR engine here is to screen out picture works containing more features (such as color features, shape features, texture features, etc.) in the picture works to be authenticated from the authenticated picture works, so as to narrow the scope of authentication comparison, save server resources, and further improve the efficiency of copyright authentication of picture works. Of course, it should be noted that, in specific implementation, the similarity value between the digital work to be authenticated and the authenticated image work in the preset authentication work library can also be calculated directly based on the SIFT feature vector, and then whether the copyright authentication is passed can be determined based on the similarity value. However, in comparison, feature extraction and similarity calculation through the CBIR engine are simpler and more efficient than extracting SIFT feature vectors. Therefore, in this embodiment, the CBIR engine is first used for initial screening, and then combined with the subsequent calculation of the similarity value based on the SIFT feature vector. Compared with the method of directly calculating the similarity value based on the SIFT feature vector, the efficiency of copyright authentication of the work can be further improved.
[0171] Then, SIFT (Scale-invariant feature transform) feature vectors of the digital work to be authenticated and similar picture works are extracted respectively, that is, the first SIFT feature vector of the digital work to be authenticated is extracted, and the second SIFT feature vector of the similar picture work is extracted. The SIFT feature vector extraction process is as follows: a series of key points are detected in the image, which are independent of scale scaling, rotation and brightness changes, and then the gradient direction values are assigned to the key points to obtain the SIFT feature vector of an image. The specific extraction process is consistent with the prior art and will not be described in detail here.
[0172] Then, the cosine distance between the first SIFT feature vector and the second SIFT feature vector is calculated to obtain a second calculation result. In this embodiment, the cosine distance is used to characterize the similarity between the two feature vectors. In a specific embodiment, other parameters may also be used to characterize, such as the Euclidean distance. Finally, it is detected whether there is a cosine distance greater than the second preset threshold in the second calculation result; wherein, the second preset threshold may be the same as or different from the first preset threshold, and may be set according to actual needs, for example, it may also be set to 0.8, and no specific limitation is made here. If there is a cosine distance greater than the second preset threshold, it is considered that there is a citation / plagiarism relationship, and the copyright certification is determined to be unsuccessful; if there is no cosine distance greater than the second preset threshold, the copyright certification is determined to be successful.
[0173] Similarly, in the specific implementation, the processes of "extracting the first SIFT feature vector of the digital work to be authenticated, and extracting the second SIFT feature vector of the similar picture work, and then calculating the cosine distance between the first SIFT feature vector and the second SIFT feature vector" and "detecting whether there is a cosine distance greater than the second preset threshold" can be performed simultaneously. Once it is detected that the cosine distance corresponding to a similar picture work is greater than the second preset threshold, it is considered that there is a citation / plagiarism relationship, and the copyright authentication can be determined to be unsuccessful. At this time, there is no need to extract SIFT feature vectors of other similar picture works and the digital work to be authenticated, nor is there any need to calculate the corresponding cosine distance and subsequent detection steps, which can save server resources and further improve the efficiency of copyright authentication.
[0174] In this embodiment, if the work type is an audio work, step S40 may further include:
[0175] Step c31, converting the digital work to be authenticated into a text work type to obtain an audio text work to be authenticated;
[0176] Step c32, calculating a third similarity value between the audio text work to be authenticated and authenticated audio text works in a preset authentication work library through a preset document search engine;
[0177] Step c33, retrieving a third preset number of similar audio text works from the authenticated audio text works according to the third similarity value;
[0178] Step c34, calculating the second longest common subsequence between the similar audio file work and the audio text work to be authenticated, and calculating the length ratio between the similar audio text work and the audio text work to be authenticated according to the length of the second longest common subsequence, to obtain a third calculation result;
[0179] Step c35, detecting whether there is a length ratio greater than a third preset threshold in the third calculation result;
[0180] Step c36, if there is a length ratio greater than the third preset threshold, it is determined that the copyright authentication fails;
[0181] Step c37, if there is no length ratio greater than the third preset threshold, it is determined that the copyright authentication passes.
[0182] In this embodiment, if the work type is an audio work, its corresponding copyright authentication process is as follows:
[0183] Convert the digital work to be authenticated into a text work type to obtain the audio text work to be authenticated. Then, calculate the third similarity value between the digital work to be authenticated and the authenticated audio text works in the preset authentication work library through a preset document search engine. Among them, the preset document search engine can optionally be an ES search engine. The authenticated audio text works in the preset authentication work library are obtained by converting audio works into text work types based on a speech recognition tool and can be stored in advance by the ES search engine according to a specific index structure for subsequent search.
[0184] After obtaining the third similarity value between the digital work to be authenticated and the authenticated audio text works, screen out the third preset number of similar text works from the authenticated audio text works according to the third similarity value. Among them, the third preset number can be the same as or different from the first preset number and the second preset number, and can be set according to actual needs, and no specific limitation is made here. When screening, sort the third similarity values in descending order and take the first third preset number of text works as the similar text works. It should be noted that the purpose of screening here is to screen out text works from the authenticated audio text works that contain all or most of the segmented words in the audio text work to be authenticated, so as to narrow the scope of authentication comparison, save server resources, and further improve the efficiency of copyright authentication.
[0185] Then, calculate the second longest common subsequence between the similar text audio work and the audio text work to be authenticated. Among them, the "second" in the second longest common subsequence has no substantial meaning and is only used to distinguish it from the above-mentioned first longest common subsequence. Furthermore, calculate the length ratio between the similar text audio work and the digital work to be authenticated according to the length of the second longest common subsequence to obtain a third calculation result, and detect whether there is a length ratio greater than the third preset threshold in the third calculation result. Among them, the third preset threshold can be the same as or different from the first preset threshold and the second preset threshold, and can be set according to actual needs. For example, it can be set to 0.8, and no specific limitation is made here. If there is a length ratio greater than the third preset threshold, it is considered that there is a citation / copying relationship, and at this time, it is determined that the copyright authentication fails; if there is no length ratio greater than the third preset threshold, it is determined that the copyright authentication passes.
[0186] It should be noted that, in specific implementation, the processes of "calculating the second longest common subsequence between similar audio text works and the digital work to be authenticated, and calculating the length ratio between the similar audio text works and the digital work to be authenticated based on the length of the second longest common subsequence" and "detecting whether there is a length ratio greater than the third preset threshold" can be performed simultaneously, that is, the second longest common subsequence between each similar audio text work and the digital work to be authenticated is calculated in turn, and then the length ratio between the similar audio text works and the digital work to be authenticated is calculated based on the length of the second longest common subsequence, and then whether the length ratio is greater than the third preset threshold. Once it is detected that the length ratio corresponding to a certain similar audio text work is greater than the third preset threshold, it is considered that there is a citation / plagiarism relationship, and the copyright authentication can be determined to be unsuccessful. At this time, there is no need to calculate the second longest common subsequence between other similar audio text works and the digital work to be authenticated and subsequent steps, which can save server resources and further improve the efficiency of copyright authentication.
[0187] Furthermore, based on the above-mentioned first, second and fourth embodiments, a fifth embodiment of the copyright authentication method of the present invention is proposed.
[0188] In this embodiment, after step S40, the copyright authentication method may further include:
[0189] Step A, when the copyright authentication is passed, obtaining the work information of the digital work to be authenticated;
[0190] In this embodiment, when the copyright authentication is passed, the work information of the digital work to be authenticated is obtained, wherein part of the work information (such as the author of the work, the type of work, etc.) can be obtained according to the digital work copyright authentication request, or a corresponding prompt window can be generated, and the work information can be obtained after the user fills in the corresponding work information based on the prompt window; the other part of the information can be generated by the copyright authentication system, such as the authentication time and the md5 (Message-Digest Algorithm) value of the work, wherein the authentication time can directly obtain the time when the copyright authentication is passed, and the md5 value of the work can be obtained through the corresponding program after obtaining the digital work to be authenticated, for example, the byte information of the file is obtained, and the md5 encryption is performed through the MessageDigest class in the second step, and the third step is converted into a hexadecimal md5 code value. The work information may include but is not limited to: the author of the work, the type of work, the authentication time and the md5 value of the digital work to be authenticated.
[0191] Step B, generating corresponding copyright authentication information based on the work information, and generating a data upload request according to the copyright authentication information;
[0192] Then, the corresponding copyright authentication information is generated based on the work information. Specifically, the work information can be converted into a preset data format, such as a json (a lightweight data exchange format) data structure, to obtain the copyright authentication information. After the copyright authentication information is generated, a data chain request is generated based on the copyright authentication information. Specifically, the sha256 (Secure Hash Algorithm 256) algorithm can be used to generate a hash value of the copyright authentication information, and then a data chain request is generated based on the hash value.
[0193] Step C, sending the data chain request to the copyright authentication alliance chain, so that the copyright authentication alliance chain completes the chain operation of the digital work to be authenticated based on the data chain request.
[0194] Finally, the data chain request is sent to the copyright certification alliance chain, so that the copyright certification alliance chain completes the chain operation of the certified digital work based on the data chain request. Among them, the copyright certification alliance chain is mainly composed of copyright certification agency nodes, notary agency nodes, judicial agency nodes and external nodes. The data chain request can be sent to the external node in the copyright certification alliance chain, and then all nodes together complete the chain operation of the certified digital work based on the general consensus algorithm, that is, the hash value in the data chain request is written into the alliance chain for permanent retention.
[0195] Of course, it is understandable that when copyright certification is passed, in addition to reporting the data to the copyright certification alliance chain, the digital work to be certified that has passed the copyright certification can also be stored in the preset certified work library to detect whether subsequent works plagiarize certified works; at the same time, a message prompt indicating the success of copyright certification is returned to the user end to inform the user.
[0196] In this embodiment, when the copyright authentication system determines that the copyright authentication has passed, it can generate a corresponding data chain request and send it to the copyright authentication alliance chain, so that the copyright authentication alliance chain can complete the chain operation based on the data chain request. Based on the unmodifiable nature of the blockchain, the protection of the copyright of digital works can be achieved. Once the unauthorized dissemination of other people's works occurs in the future, the copyright owner can file a complaint against the infringement based on the authentication information on the blockchain, reducing the difficulty of rights protection.
[0197] In the prior art, there are also copyright authentication schemes based on blockchain. For example, copyright authentication agencies, notary agencies, judicial agencies, and several self-media people act as nodes to jointly constitute a copyright authentication blockchain platform to review and authenticate the copyright of digital works. However, the review of works and copyright authentication are still completed manually offline, and only the final authentication results are written into the blockchain platform. This method does not shorten the authentication cycle and improve the efficiency of copyright authentication. At the same time, the current blockchain-based schemes mostly use public chain methods to organize blockchains to ensure the stability of system operation. In addition, self-media people or creators are required to apply for and join the blockchain platform, which increases the use cost of the creators, and others can know the operation behavior of each node, and data privacy cannot be effectively guaranteed. In this regard, the present invention also provides a copyright authentication system.
[0198] Reference Figure 3 , Figure 3 Schematic diagram of the system architecture of the copyright authentication system of the present invention.
[0199] In this embodiment, if Figure 3 As shown, the copyright authentication system includes copyright authentication equipment and a copyright authentication alliance chain; of course, it can also include a user end.
[0200] Among them, copyright authentication equipment is such as Figure 1 The copyright authentication device shown is used to execute the various steps in the above-mentioned copyright authentication method embodiment. The specific functions and implementation processes can refer to the above-mentioned embodiments and will not be repeated here.
[0201] The copyright authentication alliance chain is used to receive a data upload request sent by the copyright authentication device;
[0202] Based on the data on-chain request, the data information to be on-chain is obtained, and based on the consensus algorithm, the on-chain operation of the data information to be on-chain is completed.
[0203] In this embodiment, the copyright authentication alliance chain can be used to receive the data chain request sent by the copyright authentication device, wherein the copyright authentication alliance chain is mainly composed of the copyright authentication agency node, the notary agency node, the judicial agency node and the external node, and the data chain request sent by the copyright authentication device can be received through the external node. Then, based on the data chain request, the data information to be chained is obtained, wherein the data information to be chained can be a hash value generated based on the copyright authentication information, and then the chain operation of the hash value is completed based on the consensus algorithm, that is, the hash value is written into the alliance chain for permanent retention.
[0204] In addition, to ensure the authority and security of copyright authentication, usually only requests from copyright authentication devices are responded to. Correspondingly, when receiving a data chain request, the external node can obtain the corresponding device IP (Internet Protocol) and then detect whether the device IP is the IP of the preset copyright authentication device. If so, the data information to be chained is obtained based on the data chain request, and the chain operation of the data information to be chained is completed based on the consensus algorithm.
[0205] By constructing the above-mentioned copyright authentication system, online authentication of the copyright of digital works can be achieved through the copyright authentication device, and copyright authentication can be completed within a time complexity of seconds for different types of digital works. Compared with the manual copyright authentication in the prior art, the present invention can reduce labor costs, shorten the copyright authentication cycle, and improve the efficiency of copyright authentication, so that the copyright of the author of the work can be protected in time. In addition, in this embodiment, self-media people or creators do not need to apply for and join the blockchain platform to achieve copyright authentication and protection, thereby reducing the use cost of the creator. At the same time, others cannot obtain the creator's behavior and operation, which can ensure data privacy.
[0206] The invention also provides a copyright authentication device.
[0207] Reference Figure 4 , Figure 4 Schematic diagram of the functional modules of the first embodiment of the copyright authentication device of the present invention.
[0208] like Figure 4 As shown, the copyright authentication device includes:
[0209] The first acquisition module 10 is used to acquire the digital work to be authenticated and the type of the work according to the digital work copyright authentication request when receiving the digital work copyright authentication request;
[0210] A processing module 20 is used to determine a corresponding processing strategy and a target review classification model according to the type of the work, and to process the digital work to be authenticated based on the processing strategy to obtain a target input object;
[0211] The review module 30 is used to input the target input object into the target review classification model, obtain the review result, and determine whether the review is passed based on the review result;
[0212] The copyright authentication module 40 is used to compare the digital work to be authenticated with the authenticated digital works in the preset authentication work library to perform copyright authentication when the digital work is passed the review.
[0213] Furthermore, the processing module 20 is specifically used for:
[0214] If the type of the work is a text work, the corresponding processing strategy is determined to be the first processing strategy, and the target review classification model is determined to be the first review classification model;
[0215] Performing word segmentation processing on the digital work to be authenticated based on the first processing strategy to obtain a first word segmentation text;
[0216] Inputting the first segmented word text into a preset word vector model to obtain a first word vector for each segmented word in the first segmented word text;
[0217] A first document vector corresponding to the digital work to be authenticated is obtained according to the first word vector, wherein the target input object is the first document vector.
[0218] Furthermore, the processing module 20 is also specifically used for:
[0219] If the work type is a picture work, determining the corresponding processing strategy to be the second processing strategy, and determining the target review classification model to be the second review classification model;
[0220] The digital work to be authenticated is preprocessed based on the second processing strategy to obtain an input image, wherein the preprocessing includes scaling processing and grayscale processing, and the target input object is the input image.
[0221] Furthermore, the processing module 20 is also specifically used for:
[0222] If the work type is an audio work, determining the corresponding processing strategy to be the third processing strategy, and determining the target review classification model to be the third review classification model;
[0223] Converting the digital work to be authenticated into a text work type based on the third processing strategy to obtain a converted digital work to be authenticated;
[0224] Performing word segmentation processing on the converted digital work to be authenticated to obtain a second word segmentation text;
[0225] Inputting the second segmented word text into a preset word vector model to obtain a second word vector for each segmented word in the second segmented word text;
[0226] A second document vector corresponding to the converted digital work to be authenticated is obtained according to the second word vector, wherein the target input object is the second document vector.
[0227] Furthermore, the target review classification model includes multiple ones, the number of the review results is the same as the number of the target review classification models, and the review module 30 is specifically used for:
[0228] Check whether the multiple review results are all qualified;
[0229] If all of the above review results are qualified, the review is considered to be passed;
[0230] If at least one of the multiple review results is unqualified, the review is determined to be unsuccessful.
[0231] Furthermore, if the work type is a text work, the copyright authentication module 40 is specifically used to:
[0232] Calculating a first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library by using a preset document search engine;
[0233] Filtering a first preset number of similar text works from the authenticated text works according to the first similarity value;
[0234] Calculating a first longest common subsequence between the similar text work and the digital work to be authenticated, and calculating a length ratio between the similar text work and the digital work to be authenticated based on the length of the first longest common subsequence to obtain a first calculation result;
[0235] Detecting whether there is a length ratio greater than a first preset threshold in the first calculation result;
[0236] If there is a length ratio greater than the first preset threshold, it is determined that the copyright authentication fails;
[0237] If there is no length ratio greater than the first preset threshold, it is determined that the copyright authentication is passed.
[0238] Furthermore, the copyright authentication module 40 is also specifically used for:
[0239] Performing word segmentation processing on the digital work to be authenticated by using a preset document search engine to obtain a word segmentation set;
[0240] Performing an inverted index on the certified text works in the preset certified works library through the preset document search engine, and calculating the score corresponding to each word in the word set according to the inverted index result;
[0241] The scores of the segmented words are summed up to obtain a first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library.
[0242] Furthermore, if the work type is a picture work, the copyright authentication module 40 is further specifically used for:
[0243] Calculating a second similarity value between the digital work to be authenticated and authenticated image works in a preset authentication work library by using a preset image retrieval engine;
[0244] Filtering a second preset number of similar picture works from the authenticated picture works according to the second similarity value;
[0245] Extracting a first scale-invariant feature transform (SIFT) feature vector of the digital work to be authenticated, and extracting a second SIFT feature vector of the similar image work;
[0246] Calculating a cosine distance between the first SIFT feature vector and the second SIFT feature vector to obtain a second calculation result;
[0247] Detecting whether there is a cosine distance greater than a second preset threshold in the second calculation result;
[0248] If there is a cosine distance greater than the second preset threshold, it is determined that the copyright authentication fails;
[0249] If there is no cosine distance greater than the second preset threshold, it is determined that the copyright authentication is passed.
[0250] Furthermore, if the work type is an audio work, the copyright authentication module 40 is further specifically used for:
[0251] Converting the digital work to be authenticated into a text work type to obtain an audio text work to be authenticated;
[0252] Calculating a third similarity value between the audio text work to be authenticated and authenticated audio text works in a preset authentication work library through a preset document search engine;
[0253] Retrieving a third preset number of similar audio text works from the authenticated audio text works according to the third similarity value;
[0254] Calculating the second longest common subsequence between the similar audio file work and the audio text work to be authenticated, and calculating the length ratio between the similar audio text work and the audio text work to be authenticated according to the length of the second longest common subsequence, to obtain a third calculation result;
[0255] Detecting whether there is a length ratio greater than a third preset threshold in the third calculation result;
[0256] If there is a length ratio greater than the third preset threshold, it is determined that the copyright authentication fails;
[0257] If there is no length ratio greater than the third preset threshold, it is determined that the copyright authentication is passed.
[0258] Furthermore, the copyright authentication device further includes:
[0259] A second acquisition module is used to acquire the work information of the digital work to be authenticated when the copyright authentication is passed;
[0260] A generation module, used to generate corresponding copyright authentication information based on the work information, and generate a data upload request according to the copyright authentication information;
[0261] The sending module is used to send the data chain request to the copyright authentication alliance chain, so that the copyright authentication alliance chain completes the chain operation of the digital work to be authenticated based on the data chain request.
[0262] Among them, the functional implementation of each module in the above-mentioned copyright authentication device corresponds to the various steps in the above-mentioned copyright authentication method embodiment, and its functions and implementation processes will not be repeated here one by one.
[0263] The present invention also provides a computer-readable storage medium on which a copyright authentication program is stored. When the copyright authentication program is executed by a processor, the steps of the copyright authentication method described in any of the above embodiments are implemented.
[0264] The specific embodiments of the computer-readable storage medium of the present invention are basically the same as the embodiments of the above-mentioned copyright authentication method, and will not be described in detail here.
[0265] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0266] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0267] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0268] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A copyright authentication method, It is characterized in that The copyright authentication method comprises: Upon receiving a digital work copyright authentication request, obtaining the digital work to be authenticated and the type of the work according to the digital work copyright authentication request; Determine a corresponding processing strategy and a target review classification model according to the type of the work, and process the digital work to be authenticated based on the processing strategy to obtain a target input object; Input the target input object into the target review classification model to obtain a review result, and determine whether the review is passed based on the review result; When the review is passed, the digital work to be authenticated is compared with the authenticated digital works in the preset authentication work library to perform copyright authentication; The step of determining a corresponding processing strategy and a target review classification model according to the type of the work, and processing the digital work to be authenticated based on the processing strategy to obtain a target input object includes: If the type of the work is a text work, the corresponding processing strategy is determined to be the first processing strategy, and the target review classification model is determined to be the first review classification model; If the work type is an audio work, the corresponding processing strategy is determined to be the third processing strategy, and the target review classification model is determined to be the third review classification model, the third review classification model is the same as the first review classification model, or the third review classification model is a classification model trained based on the first review classification model, and the first review classification model is a binary classification model; Converting the digital work to be authenticated into a text work type based on the third processing strategy to obtain a converted digital work to be authenticated; Performing word segmentation processing on the converted digital work to be authenticated to obtain a second word segmentation text; Inputting the second segmented word text into a preset word vector model to obtain a second word vector for each segmented word in the second segmented word text; The second word vectors are added according to corresponding dimensions to obtain a second document vector corresponding to the converted digital work to be authenticated, wherein the target input object is the second document vector.
2. The copyright authentication method according to claim 1, It is characterized in that The step of determining a corresponding processing strategy and a target review classification model according to the work type, and processing the digital work to be authenticated based on the processing strategy to obtain a target input object comprises: If the type of the work is a text work, determining the corresponding processing strategy to be the first processing strategy, and determining the target review classification model to be the first review classification model; Performing word segmentation processing on the digital work to be authenticated based on the first processing strategy to obtain a first word segmentation text; Inputting the first segmented word text into a preset word vector model to obtain a first word vector for each segmented word in the first segmented word text; A first document vector corresponding to the digital work to be authenticated is obtained according to the first word vector, wherein the target input object is the first document vector.
3. The copyright authentication method according to claim 1, It is characterized in that The step of determining a corresponding processing strategy and a target review classification model according to the work type, and processing the digital work to be authenticated based on the processing strategy to obtain a target input object comprises: If the work type is a picture work, the corresponding processing strategy is determined to be the second processing strategy, and the target review classification model is determined to be the second review classification model, and the second review classification model is a classification model based on a convolutional neural network; The digital work to be authenticated is preprocessed based on the second processing strategy to obtain an input image, wherein the preprocessing includes scaling processing and grayscale processing, and the target input object is the input image.
4. The copyright authentication method according to any one of claims 1 to 3, It is characterized in that The target review classification models include multiple ones, the number of the review results is the same as the number of the target review classification models, and the step of judging whether the review is passed based on the review results includes: Check whether the multiple review results are all qualified; If all of the above review results are qualified, the review is considered to be passed; If at least one of the multiple review results is unqualified, the review is determined to be unsuccessful.
5. The copyright authentication method according to claim 1, It is characterized in that If the work type is a text work, the step of comparing the digital work to be authenticated with authenticated digital works in a preset authentication work library to perform copyright authentication includes: Calculating a first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library by using a preset document search engine; Filtering a first preset number of similar text works from the authenticated text works according to the first similarity value; Calculating a first longest common subsequence between the similar text work and the digital work to be authenticated, and calculating a length ratio between the similar text work and the digital work to be authenticated based on the length of the first longest common subsequence to obtain a first calculation result; Detecting whether there is a length ratio greater than a first preset threshold in the first calculation result; If there is a length ratio greater than the first preset threshold, it is determined that the copyright authentication fails; If there is no length ratio greater than the first preset threshold, it is determined that the copyright authentication is passed.
6. The copyright authentication method as claimed in claim 5, It is characterized in that The step of calculating the first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library by using a preset document search engine comprises: Performing word segmentation processing on the digital work to be authenticated by using a preset document search engine to obtain a word segmentation set; Performing an inverted index on the certified text works in the preset certified works library through the preset document search engine, and calculating the score corresponding to each word in the word set according to the inverted index result; The scores of the segmented words are summed up to obtain a first similarity value between the digital work to be authenticated and authenticated text works in a preset authentication work library.
7. The copyright authentication method according to claim 1, It is characterized in that If the work type is a picture work, the step of comparing the digital work to be authenticated with authenticated digital works in a preset authentication work library to perform copyright authentication includes: Calculating a second similarity value between the digital work to be authenticated and authenticated image works in a preset authentication work library by using a preset image retrieval engine; Filtering a second preset number of similar picture works from the authenticated picture works according to the second similarity value; Extracting a first scale-invariant feature transform (SIFT) feature vector of the digital work to be authenticated, and extracting a second scale-invariant feature transform (SIFT) feature vector of the similar image work; Calculating a cosine distance between the first scale invariant feature transform SIFT feature vector and the second scale invariant feature transform SIFT feature vector to obtain a second calculation result; Detecting whether there is a cosine distance greater than a second preset threshold in the second calculation result; If there is a cosine distance greater than the second preset threshold, it is determined that the copyright authentication fails; If there is no cosine distance greater than the second preset threshold, it is determined that the copyright authentication is passed.
8. The copyright authentication method according to claim 1, It is characterized in that If the work type is an audio work, the step of comparing the digital work to be authenticated with authenticated digital works in a preset authentication work library to perform copyright authentication includes: Converting the digital work to be authenticated into a text work type to obtain an audio text work to be authenticated; Calculating a third similarity value between the audio text work to be authenticated and authenticated audio text works in a preset authentication work library through a preset document search engine; Retrieving a third preset number of similar audio text works from the authenticated audio text works according to the third similarity value; Calculating the second longest common subsequence between the similar audio file work and the audio text work to be authenticated, and calculating the length ratio between the similar audio text work and the audio text work to be authenticated according to the length of the second longest common subsequence, to obtain a third calculation result; Detecting whether there is a length ratio greater than a third preset threshold in the third calculation result; If there is a length ratio greater than the third preset threshold, it is determined that the copyright authentication fails; If there is no length ratio greater than the third preset threshold, it is determined that the copyright authentication is passed.
9. The copyright authentication method according to any one of claims 1 to 3 and 5 to 8, It is characterized in that The copyright authentication method further includes: When the copyright authentication is passed, obtaining the work information of the digital work to be authenticated; Generate corresponding copyright authentication information based on the work information, and generate a data upload request based on the copyright authentication information; The data chain upload request is sent to the copyright authentication alliance chain, so that the copyright authentication alliance chain completes the chain upload operation of the digital work to be authenticated based on the data chain upload request.
10. A copyright authentication device, It is characterized in that The copyright authentication device comprises: A first acquisition module is used to, upon receiving a digital work copyright authentication request, acquire the digital work to be authenticated and the work type according to the digital work copyright authentication request; A processing module, used to determine a corresponding processing strategy and a target review classification model according to the type of the work, and process the digital work to be authenticated based on the processing strategy to obtain a target input object; A review module, used for inputting the target input object into the target review classification model, obtaining a review result, and judging whether the review is passed based on the review result; A copyright authentication module is used to compare the digital work to be authenticated with the authenticated digital works in a preset authentication work library to perform copyright authentication when the digital work is passed the review; The step of determining a corresponding processing strategy and a target review classification model according to the type of the work, and processing the digital work to be authenticated based on the processing strategy to obtain a target input object includes: If the type of the work is a text work, the corresponding processing strategy is determined to be the first processing strategy, and the target review classification model is determined to be the first review classification model; If the work type is an audio work, the corresponding processing strategy is determined to be the third processing strategy, and the target review classification model is determined to be the third review classification model, the third review classification model is the same as the first review classification model, or the third review classification model is a classification model trained based on the first review classification model, and the first review classification model is a binary classification model; Converting the digital work to be authenticated into a text work type based on the third processing strategy to obtain a converted digital work to be authenticated; Performing word segmentation processing on the converted digital work to be authenticated to obtain a second word segmentation text; Inputting the second segmented word text into a preset word vector model to obtain a second word vector for each segmented word in the second segmented word text; The second word vectors are added according to corresponding dimensions to obtain a second document vector corresponding to the converted digital work to be authenticated, wherein the target input object is the second document vector.
11. A copyright authentication device, It is characterized in that The copyright authentication device comprises: a memory, a processor, and a copyright authentication program stored in the memory and executable on the processor, wherein the copyright authentication program implements the steps of the copyright authentication method as claimed in any one of claims 1 to 9 when executed by the processor.
12. A copyright certification system, It is characterized in that The copyright authentication system includes a copyright authentication device and a copyright authentication alliance chain; wherein, The copyright authentication device is the copyright authentication device as claimed in claim 11; The copyright authentication alliance chain is used to receive a data upload request sent by the copyright authentication device; Based on the data on-chain request, the data information to be on-chain is obtained, and based on the consensus algorithm, the on-chain operation of the data information to be on-chain is completed.
13. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a copyright authentication program, which, when executed by a processor, implements the steps of the copyright authentication method according to any one of claims 1 to 9.
Citation Information
Patent Citations
A text similarity analysis method and system for copyright authentication
CN109145529A
Copyright registration method and device based on block chain and terminal equipment
CN109684786A
Block chain network digital work registration method and client
CN110188515A