Multi-modal large language model-based copyright identification method and apparatus, and electronic device

Through the copyright identification method based on the multimodal large language model, the problems of large resource consumption and poor recognition effect in the existing technology are solved, and efficient identification and cross-modal copyright search for all types of copyright infringement are achieved.

CN120196960APending Publication Date: 2025-06-24NANJING XIYIN ECOMMERCE CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510341535.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art requires training a model for each type of copyright infringement identification, resulting in poor recognition of copyright types with large resource consumption and little training data.

Method used

The copyright identification method based on the multimodal large language model is adopted to realize the identification of all types of copyright infringement through an identification model. The method includes obtaining target case data and text description information, performing feature extraction, and matching with the copyright index library to determine whether there is any infringement.

Benefits of technology

It realizes the identification of all types of copyright infringement through one model, avoiding the problem of large resource consumption and poor recognition effect of copyright type based on few training data, and also supports cross-modal copyright search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196960A_ABST
    Figure CN120196960A_ABST
Patent Text Reader

Abstract

The invention discloses a copyright recognition method and device based on a multi-modal large language model and electronic equipment. The method comprises the steps of obtaining target case data and / or target text description information to be subjected to copyright identification; inputting the target case data into a trained target multi-modal large language model for feature extraction to obtain target case features of the target case data; and / or inputting the target character description information into a target multi-modal large language model for feature extraction to obtain target character description features of the target character description information; inputting the target case features into a created copyright index database for matching to obtain a first matching result; and / or, inputting the target character description feature into a copyright index database for matching to obtain a second matching result; whether the target case data is infringement or not is determined based on the first matching result and / or the second matching result, and the recognition effect of copyright types with few training data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of artificial intelligence. More specifically, this disclosure relates to a copyright recognition method, apparatus, and electronic device based on a multimodal large language model. Background Art

[0002] Intellectual Property (IP) protection refers to protecting people's rights and interests in innovation, original works, and other intellectual achievements through legal means. With the development of globalization, cross-border technologies and products are becoming increasingly frequent in circulation, and IP infringement incidents occur frequently. Therefore, the demand for IP infringement recognition is becoming increasingly prominent.

[0003] Currently, conventional IP infringement recognition methods usually involve training an infringement recognition model for each type of copyright, and then recognizing the infringement of that type of copyright based on the trained infringement recognition model. However, the training data for different types of copyrights are different. For copyrights with less training data, the recognition effect of the trained infringement recognition model is poor. Moreover, training an infringement recognition model for each type of copyright will also result in relatively large resource consumption.

[0004] In view of this, there is an urgent need to provide a copyright recognition method, apparatus, and electronic device based on a multimodal large language model, so as to enable the recognition of all types of copyright infringement through a single trained recognition model, avoiding the problems of large resource consumption caused by training multiple infringement recognition models and poor recognition effects of infringement recognition models trained based on copyright types with less training data. Summary of the Invention

[0005] To at least solve one or more of the above-mentioned technical problems, this disclosure proposes a copyright recognition method, apparatus, and electronic device based on a multimodal large language model in multiple aspects.

[0006] In a first aspect, the present disclosure provides a copyright recognition method based on a multimodal large language model, the method comprising: obtaining target case data to be recognized for copyright and / or target text description information of the target case data; the target case data being video data, audio data, image data or text data; inputting the target case data into a trained target multimodal large language model for feature extraction to obtain target case features of the target case data; and / or inputting the target text description information into the target multimodal large language model for feature extraction to obtain target text description features of the target text description information, wherein the target multimodal large language model is trained based on training data of multiple elements that need to be protected by copyright, the elements being the smallest units that need to be protected by copyright; matching the target case features with each piece of data in a pre-created copyright index library to obtain a first matching result; and / or matching the target text description features with each piece of data in the copyright index library to obtain a second matching result; the copyright index library comprising: multiple pieces of data, each piece of data including corresponding element features and text description features; displaying the first matching result and / or the second matching result to determine whether the target case data is infringing based on the displayed first matching result and / or second matching result.

[0007] In some embodiments, the target multimodal large language model is trained through the following steps: for each element that needs to be protected by copyright, obtaining multiple pieces of training data containing the element and text description information of the element; storing the first associated data pair and the second associated data pair in an associated manner; wherein the first associated data pair is a data pair composed of an element and text description information of the element; the second associated data pair is a data pair composed of training data and text description information of the element in the training data; inputting the first associated data pair and / or the second associated data pair into an initial multimodal large language model for training to obtain the target multimodal large language model.

[0008] In some embodiments, the copyright index library is obtained through the following steps: inputting the elements in the second associated data pair into the target multimodal large language model for feature extraction to obtain element features; and inputting the text description information in the second associated data pair into the target multimodal large language model for feature extraction to obtain text description features; storing the element features and the text description features in an associated manner to form the copyright index library.

[0009] In some embodiments, the elements that need to be protected by copyright include images, text, audio, and video.

[0010] In some embodiments, matching the target case features with each piece of data in the created copyright index library to obtain a first matching result includes: calculating the similarity between the target case features and the element features in each piece of data in the copyright index library; taking the element features corresponding to the top N similarities in descending order as the first matching result; where N is a preset positive integer; and / or, matching the target text description features with each piece of data in the copyright index library to obtain a second matching result includes: calculating the similarity between the target text description features and the text description features in each piece of data in the copyright index library; taking the text description features corresponding to the top M similarities in descending order as the second matching result; where M is a preset positive integer.

[0011] In some embodiments, matching the target case features with each piece of data in the created copyright index library to obtain a first matching result includes: calculating the similarity between the target case features and the element features in each piece of data in the copyright index library; taking the element features corresponding to the similarities greater than or equal to a first set threshold as the first matching result; and / or, inputting the target text description features into the copyright index library for matching to obtain a second matching result includes: calculating the similarity between the target text description features and the text description features in each piece of data in the copyright index library; taking the element features corresponding to the similarities greater than or equal to a second set threshold as the second matching result.

[0012] In some embodiments, the similarity is cosine similarity or Euclidean similarity.

[0013] In a second aspect, the present disclosure provides a copyright recognition device based on a multimodal large language model, the device comprising: an information acquisition module for acquiring target case data to be subject to copyright recognition and / or target text description information of the target case data; the target case data being video data, audio data, image data or text data; a feature extraction module for inputting the target case data into a trained target multimodal large language model for feature extraction to obtain target case features of the target case data; and / or inputting the target text description information into the target multimodal large language model for feature extraction to obtain target text description features of the target text description information, wherein the target multimodal large language model is trained based on training data of a plurality of elements that need to be protected by copyright, the elements being the smallest units that need to be protected by copyright; a matching module for inputting the target case features into a created copyright index library for matching to obtain a first matching result; and / or inputting the target text description features into the copyright index library for matching to obtain a second matching result; the copyright index library comprising: element features and text description features in one-to-one correspondence; an infringement determination module for displaying the first matching result and / or the second matching result to determine whether the target case data is infringing based on the displayed first matching result and / or second matching result.

[0014] In a third aspect, the present disclosure provides an electronic device comprising: a processor configured to execute program instructions; and a memory configured to store the program instructions, which, when loaded and executed by the processor, cause the processor to execute the copyright recognition method based on a multimodal large language model as described in the first aspect or any optional embodiment of the first aspect.

[0015] In a fourth aspect, the present disclosure provides a computer-readable storage medium storing program instructions, which, when loaded and executed by a processor, cause the processor to execute the copyright recognition method based on a multimodal large language model as described in the first aspect or any optional embodiment of the first aspect.

[0016] Through a copyright recognition method, device, and electronic device based on a multimodal large language model provided as above, in the embodiments of this disclosure, the target case data to be copyright-recognized and / or the target text description information of the target case data are input into a trained target multimodal large language model for feature extraction to obtain target case features and / or target text description features, and based on the target case features and / or target text description features, a match is made in a created copyright index library to obtain a first match result and / or a second match result, and based on the first match result and / or the second match result, it is determined whether the target case data is infringing. This realizes the recognition of copyright infringement of all types (i.e., videos, audios, images, texts, etc.) through a trained recognition model (i.e., the target multimodal large language model), avoiding the problem of high resource consumption caused by training multiple infringement recognition models and the problem of poor recognition effect of infringement recognition models trained based on copyright types with little training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of this disclosure will become readily understood. In the drawings, several embodiments of this disclosure are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where:

[0018] Figure 1 Shows an exemplary flowchart of a copyright recognition method based on a multimodal large language model according to some embodiments of this disclosure;

[0019] Figure 2 Shows a structural diagram of a copyright recognition system of a multimodal large language model according to some embodiments of this disclosure;

[0020] Figure 3 Shows an exemplary structural block diagram of a copyright recognition device based on a multimodal large language model according to some embodiments of this disclosure;

[0021] Figure 4 Shows an exemplary structural block diagram of an electronic device according to some embodiments of this disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Next, the technical solutions in the embodiments of this disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of this disclosure. Based on the embodiments in this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this disclosure.

[0023] It should be understood that the terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0024] It should also be understood that the terms used in this disclosure specification are for the purpose of describing particular embodiments only and are not intended to limit this disclosure. As used in this disclosure specification and claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should further be understood that the term "and / or" as used in this disclosure specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0025] As used in this specification and claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0026] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.

[0027] Exemplary Application Scenarios

[0028] Intellectual property protection refers to protecting people's rights and interests in innovations, original works, and other intellectual achievements through legal means. The current methods for identifying intellectual property infringement are generally as follows: for each type of copyright, an infringement identification model is trained, and then the infringement of this type of copyright is identified based on the trained infringement identification model. That is, the current intellectual property infringement identification can only achieve image search by image, text search by text, language search by language, and video search by video. However, the training data for different types of copyrights are different. For copyrights with less training data, the identification effect of the trained infringement identification model is poor. And training an infringement identification model for each type of copyright will also result in relatively large resource consumption.

[0029] In view of this, the embodiments of the present disclosure provide a copyright recognition solution based on a multimodal large language model, so as to realize the recognition of all types of copyright infringement through one recognition model (i.e., the multimodal large language model), avoiding the problem of high resource consumption caused by training multiple infringement recognition models (i.e., videos, audios, images, texts, etc.), reducing the system development cost, and improving the model training efficiency; at the same time, it also solves the problem of poor recognition effect of the infringement recognition model trained based on copyright types with less training data, and reduces the data acquisition cost.

[0030] Figure 1 FIG. shows an exemplary flowchart of a copyright recognition method 100 based on a multimodal large language model according to some embodiments of the present disclosure. It can be understood that the copyright recognition method 100 of the multimodal large language model can be executed by any suitable device with data processing capabilities, such as but not limited to terminal devices, processors, and servers, etc.

[0031] As Figure 1 shown, the copyright recognition method 100 based on the multimodal large language model includes: Step S110: Obtain target case data to be recognized for copyright and / or target text description information of the target case data; Step S120: Input the target case data into a trained target multimodal large language model for feature extraction to obtain target case features of the target case data; and / or, input the target text description information into the target multimodal large language model for feature extraction to obtain target text description features of the target text description information; Step S130: Match the target case features with each piece of data in the created copyright index library to obtain a first matching result; and / or, match the target text description features with each piece of data in the copyright index library to obtain a second matching result; The copyright index library includes: multiple pieces of data, and each piece of data includes corresponding element features and text description features; Step S140: Display the first matching result and / or the second matching result to determine whether the target case data is infringed based on the displayed first matching result and / or the second matching result.

[0032] Exemplarily, in the embodiments of the present disclosure, the target case data in the above step S110 refers to the specific data that needs to be recognized for copyright, such as video data, audio data, image data, or text data. Among them, video data is dynamic media data composed of a series of consecutive image frames and audio; audio data can include music, language, etc.; image data is static visual information, such as photos, paintings, image trademarks, etc.; text data refers to the information composed of characters, including articles, books, etc.

[0033] Based on the above description, the target case data can be video data downloaded from a video platform or video data collected by video capture devices such as cameras, audio data obtained from an audio file library or audio data collected by sound pickup components such as microphones, text data read from a text file or text data copied from document editing software such as Word or WPS, or image data obtained from an image database or image data collected by image capture devices such as scanners or digital cameras. Of course, the above target case data can also be manually input or obtained by other reasonable acquisition methods, and the embodiments of the present disclosure do not make specific limitations on this.

[0034] In the embodiments of the present disclosure, the target text description information in step S110 above refers to the text description information for the target case data, which can provide more detailed description information for the copyright identification of the target case data and can be used to assist in copyright identification. In the embodiments of the present disclosure, the target text description information can be manually input by a person based on the target case data.

[0035] For example, when the target case data is video data, the target text description information can be a detailed text description written by a person after watching the video according to the video content, which can include key plots, scenes, character actions, etc. involved in the video data; when the target case data is audio data, the target text description information can be a detailed text description written by a person after listening to the audio according to the audio content, for example, it can include the lyrics of a song, the theme and main content of a speech, etc. When the target case data is image data, text description information can be written around the theme, color, composition, main elements, etc. of the image. When the target case data is text data, text description information can be written around the core content, creation background, key viewpoints, etc. of the text; or for each type of case data, a text description template can be set in advance, and then based on the obtained target case data, the target text description information can be formed according to its corresponding text description template.

[0036] As a specific implementation manner of the present disclosure, when the target case data is a shoe with a Nike logo, the above target description information can be a Nike logo, the font is Futura, and the placement is horizontal.

[0037] Exemplarily, in the embodiments of the present disclosure, the multi-modal large language model in step S120 above is trained in advance based on the training data of various elements that need to be protected by copyright, and it is a language model that can extract the feature information of various different modal data (such as text, image, audio, video, etc.). As for the training process of the multi-modal large language model, it is described by way of example in the following embodiments and will not be elaborated here.

[0038] In the embodiments of the present disclosure, in step S120, the target case data is input into the target multimodal large language model for feature extraction to obtain the target case features of the target case data. Here, the target case features refer to a set of features that can characterize the target case data, and the target case features can reflect information such as the content, style, and structure of the target case data. The target text description information is input into the target multimodal large language model for feature extraction to obtain the target text description features of the target text description information. Here, the target text description features refer to a set of features that can represent the target text description information, and the target text description features can reflect the semantics, grammar, emotion, etc. of the text description information.

[0039] In the embodiments of the present disclosure, the above-mentioned copyright index library refers to a database that stores element features and text description features. The copyright index library includes multiple pieces of data, and each piece of data includes corresponding element features and text description features. Here, an element refers to the smallest unit that needs copyright protection, which can be, for example, a unique pattern in an image, a specific sentence in a paragraph of text, a specific melody in a piece of audio, and so on.

[0040] As a specific implementation manner of the present disclosure, the element is the Nike swoosh, and the target case data can be an image of a T-shirt with the swoosh printed on it.

[0041] Regarding the construction process of the index library, it is described by way of example in the following embodiments and will not be elaborated here for the time being.

[0042] In the embodiments of the present disclosure, there are many methods for inputting the target case features into the created copyright index library for matching to obtain the first matching result and inputting the target text description features into the copyright index library for matching to obtain the second matching result. For example, it can be implemented based on similarity ranking, or for another example, it can be implemented based on a set threshold. The embodiments of the present disclosure do not make specific limitations on this.

[0043] As an optional embodiment of the present disclosure, for the method implemented based on similarity ranking, matching the target case features with each piece of data in the created copyright index library to obtain the first matching result includes: calculating the similarity between the target case features and the element features of each piece of data in the copyright index library; taking the element features corresponding to the top N similarities in descending order as the first matching result; where N is a preset positive integer; and / or, matching the target text description features with each piece of data in the copyright index library to obtain the second matching result includes: calculating the similarity between the target text description features and the text description features of each piece of data in the copyright index library; taking the text description features corresponding to the top M similarities in descending order as the second matching result.

[0044] Exemplarily, in the embodiments of the present disclosure, the above similarity may be cosine similarity or Euclidean similarity, and the embodiments of the present disclosure do not make specific limitations thereto.

[0045] In the embodiments of the present disclosure, when actually determining the first matching result, calculate the similarity between the element features in each piece of data in the copyright index library and the target case features. After the calculation is completed, arrange all the similarity values in descending order, and select the element features corresponding to the top N similarities as the first matching result. Here, N is a preset positive integer, for example, 5, 10, and the embodiments of the present disclosure do not make specific limitations thereto. For example, when N is 5, select the element features with the top 5 similarities as the first matching result.

[0046] In the embodiments of the present disclosure, when actually determining the second matching result, calculate the similarity between the text description features in each piece of data in the copyright index library and the target text description features. After the calculation is completed, arrange all the similarity values in descending order, and select the element features corresponding to the top N similarities as the second matching result. Here, M is a preset positive integer, for example, 5, 10, and the embodiments of the present disclosure do not make specific limitations thereto. For example, when N is 5, select the text description features with the top 5 similarities as the second matching result.

[0047] It should be noted that the above M and N may be the same or different, and the embodiments of the present disclosure do not make specific limitations thereto.

[0048] As an optional embodiment of the present disclosure, for the method implemented based on a set threshold, match the target case features with each piece of data in the created copyright index library to obtain the first matching result, including: calculating the similarity between the target case features and the element features in each piece of data in the copyright index library; taking the element features corresponding to the similarity greater than or equal to the first set threshold as the first matching result; and / or, match the target text description features with each piece of data in the copyright index library to obtain the second matching result, including: calculating the similarity between the target text description features and the text description features in each piece of data in the copyright index library; taking the element features corresponding to the similarity greater than or equal to the second set threshold as the second matching result.

[0049] Exemplarily, in the embodiments of the present disclosure, the above first set threshold may be set in advance, for example, it may be 0.8, 0.9, etc.; the above second set threshold is also set in advance, for example, it may be 0.8, 0.9, etc. It should be noted that the above first set threshold and the second set threshold may be set to be the same or different, and the embodiments of the present disclosure do not make specific limitations thereto.

[0050] In the embodiments of the present disclosure, when actually determining the first matching result, the similarity between the element features in each piece of data in the copyright index library and the target case features is calculated. After the calculation is completed, for each similarity, it is compared with the first set threshold, and the element features corresponding to the similarity greater than or equal to the first set threshold are used as the first matching result. For example, if the first set threshold is 0.8, and the similarity between a certain calculated element feature and the target case feature is 0.85, then this element feature is included in the first matching result.

[0051] In the embodiments of the present disclosure, when actually determining the second matching result, the similarity between the text description features in each piece of data in the copyright index library and the target text description features is calculated. After the calculation is completed, for each similarity, it is compared with the second set threshold, and the text description features corresponding to the similarity greater than or equal to the second set threshold are used as the second matching result. For example, if the second set threshold is 0.8, and the similarity between a certain calculated text description feature and the target text description feature is 0.85, then this text description feature is included in the second matching result.

[0052] Exemplarily, in the embodiments of the present disclosure, the matched first matching result and / or the second matching result can be displayed on a front-end display device (such as a computer display screen, a mobile phone display screen, etc.), and the user can intuitively see the element features and / or text description features with a relatively high similarity to the target case data and / or the target text description information.

[0053] In the embodiments of the present disclosure, there are many methods for determining whether the target case data is infringing based on the first matching result and / or the second matching result. For example, if there are element features or text description features that are highly similar to the known copyright-protected ones in the first matching result and the second matching result (for example, the similarity is greater than the third set threshold (for example, 0.9)), then it can be preliminarily determined that the target case data may be infringing; if there is no such situation in both the first matching result and the second matching result, then it can be preliminarily determined that the target case data is not infringing. Of course, the final infringement judgment still requires the user to make a judgment in combination with relevant laws, regulations, industry standards, etc.

[0054] In the embodiments of the present disclosure, the obtained target case data to be copyright-identified and / or the target text description information of the target case data are input into a trained target multi-modal large language model for feature extraction to obtain target case features and / or target text description features, and based on the target case features and / or target text description features, a match is made in a pre-created copyright index library to obtain a first matching result and / or a second matching result, and based on the first matching result and / or the second matching result, it is determined whether the target case data infringes copyright. This realizes the identification of copyright infringement of all types (i.e., videos, audios, images, texts, etc.) through a trained recognition model (i.e., the target multi-modal large language model), avoiding the problem of high resource consumption caused by training multiple infringement recognition models and the problem of poor recognition effect of the infringement recognition model trained based on copyright types with less training data. And in the embodiments of the present disclosure, through the multi-modal large language model, image-to-text and text-to-image searches can be realized. For example, when only the target case data (image data) is input into the target large language model, image-to-text search can be realized; for another example, when only the target description text information is input into the target large language model, text-to-image search can be realized, achieving cross-modal copyright search.

[0055] As an optional embodiment of the present disclosure, the above-mentioned target multi-modal large language model is trained through the following steps: for each element that needs copyright protection, multiple pieces of training data containing the element and the text description information of the element are obtained; the first associated data pair and the second associated data pair are stored in an associated manner; wherein, the first associated data pair is a data pair composed of the element and the text description information of the element; the second associated data pair is a data pair composed of the training data and the text description information of the element in the training data; the first associated data pair and / or the second associated data pair are input into an initial multi-modal large language model for training to obtain the target multi-modal large language model.

[0056] Exemplarily, in the embodiments of the present disclosure, the elements to be protected include images, texts, audios, and videos. For each element to be protected, first, determine the type of the element that needs copyright protection. If it is an image element, the multiple pieces of training data containing the element can be different pictures containing the image element. At the same time, write detailed text description information for the image element, describing features such as the shape, color, texture, and position in the image of the image element. If the element that needs copyright protection is a text element, the multiple pieces of training data obtained containing the text element can be article paragraphs, book chapters, etc. containing the text element. And conduct a text description of the meaning of the text element, its role in the text, and its relationship with other texts. If the element that needs copyright protection is an audio element, the multiple pieces of training data obtained containing the audio element can be music segments, voice segments, etc. containing the audio element. At the same time, conduct a text description of the frequency range, timbre characteristics, rhythm rules, etc. of the audio element. If the element that needs copyright protection is a video element, the multiple pieces of training data obtained containing the video element can be video segments containing the video element. Conduct a text description of the plot content, character characteristics, scene settings, etc. of the video element.

[0057] In the embodiments of the present disclosure, after obtaining the multiple pieces of training data containing the element and the text description information of the element, the first associated data pair and the second associated data pair can be associated and stored. Here, the first associated data pair refers to the data pair composed of the element and the text description information of the element, and the second associated data pair refers to the data pair composed of the training data and the text description information of the element in the training data.

[0058] For example, for an image element, the first associated data pair is the data pair composed of the image element and its text description information. Suppose the image element is a specific cartoon character image, and its text description information is "a cartoon character image with red hair, blue eyes, and wearing a yellow dress", and these two are combined into the first associated data pair and stored in the database.

[0059] In the embodiments of the present disclosure, when training the model, the first associated data pair and / or the second associated data pair are input into the initial multi-modal large language model for training. During the training process, the initial multi-modal large language model continuously adjusts the parameters of the model so that the model can better perform feature extraction and association analysis on the input data. For example, when inputting a picture containing a cartoon character image and the corresponding text description information, the model will learn how to extract features corresponding to the text description from the picture, and how to generate a picture feature representation that matches the text description. After a large amount of data training, the target multi-modal large language model is obtained.

[0060] As an alternative embodiment of the present disclosure, the above copyright index library is obtained through the following steps: inputting the elements in the second association data pair into a target multimodal large language model for feature extraction to obtain element features; and inputting the text description information in the second association data pair into the target multimodal large language model for feature extraction to obtain text description features; associatively storing the element features and the text description features to form a copyright index library.

[0061] Exemplarily, in the embodiment of the present disclosure, when constructing the copyright index library, only the elements (without training data) in the second association data pair need to be input into the target multimodal large language model for feature extraction to obtain element features. For example, for a picture containing a cartoon character image, the element is the cartoon character image, then the area of the cartoon character image is input into the target multimodal large language model, and the model outputs the shape feature vector, color feature vector, texture feature vector, etc. of the cartoon character image, and these feature vectors are combined into element features.

[0062] At the same time, the text description information in the second association data pair is input into the target multimodal large language model for feature extraction to obtain text description features. Still taking the above element as a cartoon character image as an example, for the text description information of the above cartoon character image, the target multimodal large language model will extract keyword feature vectors, such as the feature vectors corresponding to keywords such as "red hair", "blue eyes", "yellow dress", etc., and these feature vectors are combined into text description features.

[0063] In the embodiment of the present disclosure, after obtaining the above element features and text description features, the element features and the text description features are associatively stored to form a copyright index library for indexing during subsequent copyright identification.

[0064] Figure 2 The structure diagram of the copyright identification system of the multimodal large language model showing some embodiments of the present disclosure;

[0065] As Figure 2 shown, the copyright identification system of the multimodal large language model includes an indexing system, a retrieval system, and an auditing system. Among them, all systems are provided with indexing features by the copyright index library. The retrieval system deploys the target multimodal large language model using high-performance computing devices (such as GPU servers). Taking pictures as an example, the pictures that need to be identified for copyright infringement are input into the retrieval system. The pictures pass through the target multimodal large language model to obtain target case features, and then read the indexing features from the indexing system for retrieval, and output the matching results to the auditing system for display, and it is determined by humans based on the results displayed on the front end whether the pictures are infringing.

[0066] Figure 3The figure shows a schematic diagram of a copyright recognition device based on a multimodal large language model according to some embodiments of the present disclosure.

[0067] As Figure 3 shown, the copyright recognition device 300 based on the multimodal large language model includes: an information acquisition module 310, configured to acquire target case data to be recognized for copyright and / or target text description information of the target case data; the target case data being video data, audio data, image data or text data; a feature extraction module 320, configured to input the target case data into a trained target multimodal large language model for feature extraction to obtain target case features of the target case data; and / or, input the target text description information into the target multimodal large language model for feature extraction to obtain target text description features of the target text description information, wherein the target multimodal large language model is trained based on training data of multiple elements that need to be protected by copyright, and the element is the smallest unit that needs to be protected by copyright; a matching module 330, configured to match the target case features with each piece of data in the created copyright index library to obtain a first matching result; and / or, match the target text description features with each piece of data in the copyright index library to obtain a second matching result; the copyright index library includes: multiple pieces of data, each piece of data including corresponding element features and text description features; an infringement determination module 340, configured to display the first matching result and / or the second matching result to determine whether the target case data is infringing based on the displayed first matching result and / or second matching result.

[0068] As an optional embodiment of the present disclosure, the target multimodal large language model is obtained through module training: a training data acquisition module, configured to, for each element that needs to be protected by copyright, acquire multiple pieces of training data including the element and text description information of the element; an associated storage module, configured to store the first associated data pair and the second associated data pair in an associated manner; wherein, the first associated data pair is a data pair composed of the element and the text description information of the element; the second associated data pair is a data pair composed of the training data and the text description information of the element in the training data; a training module, configured to input the first associated data pair and / or the second associated data pair into an initial multimodal large language model for training to obtain the target multimodal large language model.

[0069] As an optional embodiment of the present disclosure, the copyright index library is obtained through the following modules: a feature acquisition module, configured to input the elements in the second associated data pair into the target multimodal large language model for feature extraction to obtain element features; and input the text description information in the second associated data pair into the target multimodal large language model for feature extraction to obtain text description features; a copyright index library acquisition module, configured to store the element features and the text description features in an associated manner to form a copyright index library.

[0070] As an alternative embodiment of the present disclosure, the elements that require copyright protection include images, texts, audio, and videos.

[0071] As an alternative embodiment of the present disclosure, the matching module 330 is specifically configured to: calculate the similarity between the target case features and the element features in each piece of data in the copyright index library; take the element features corresponding to the top N similarities in descending order as the first matching result; where N is a preset positive integer; and / or, calculate the similarity between the target text description features and the text description features in each piece of data in the copyright index library; take the text description features corresponding to the top M similarities in descending order as the second matching result; where M is a preset positive integer.

[0072] As an alternative embodiment of the present disclosure, the matching module 330 is specifically configured to: calculate the similarity between the target case features and the element features in each piece of data in the copyright index library; take the element features corresponding to the similarities greater than or equal to the first set threshold as the first matching result; and / or, calculate the similarity between the target text description features and the text description features in each piece of data in the copyright index library; take the element features corresponding to the similarities greater than or equal to the second set threshold as the second matching result.

[0073] As an alternative embodiment of the present disclosure, the similarity is cosine similarity or Euclidean similarity.

[0074] For the specific implementation manners and technical effects, reference may be made to the specific description of the embodiment of the above-mentioned copyright recognition method 100 based on the multi-modal large language model, which will not be elaborated here.

[0075] So far, the description of the Figure 3 shown device is completed.

[0076] Correspondingly, the embodiments of the present disclosure also provide Figure 4 the hardware structure diagram of the shown device, specifically as Figure 4 shown. The electronic device 400 may be the device for implementing the above method 100. As Figure 4 shown, the electronic device 400 includes: a processor 410 and a memory 420. Among them, the memory 420 is configured to store program instructions; the processor 410 is configured to load and execute the program instructions stored in the memory 420 to implement the corresponding embodiment of the copyright recognition method based on the multi-modal large language model as shown above.

[0077] As an example, the memory 420 can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as program instructions, data, and so on. For example, the memory 420 can be: volatile memory, non-volatile memory, or similar storage media. Specifically, the memory 420 can be RAM (Random Access Memory), flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as an optical disk, DVD, etc.), or similar storage media, or a combination thereof.

[0078] Thus far, the description of Figure 4 the electronic device shown is completed.

[0079] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many changes, variations, and alternative forms may be contemplated by those skilled in the art without departing from the spirit and scope of the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed in practicing the present disclosure. The appended claims are intended to define the scope of the present disclosure and thus cover equivalents or alternatives within the scope of these claims.

Claims

1. A copyright identification method based on a multimodal large language model, characterized in that: The method comprises: Acquire target case data to be subjected to copyright identification and / or target text description information of the target case data; the target case data is video data, audio data, image data or text data; Inputting the target case data into a trained target multimodal large language model for feature extraction to obtain target case features of the target case data; and / or inputting the target text description information into the target multimodal large language model for feature extraction to obtain target text description features of the target text description information, wherein the target multimodal large language model is trained based on training data of multiple elements that need to be protected by copyright, and the element is the smallest unit that needs to be protected by copyright; Matching the target case feature with each piece of data in the created copyright index library to obtain a first matching result; and / or, inputting the target text description feature into each piece of data in the copyright index library for matching to obtain a second matching result; the copyright index library includes: multiple pieces of data, each piece of data includes one-to-one corresponding element features and text description features; The first matching result and / or the second matching result are displayed to determine whether the target case data is infringing based on the displayed first matching result and / or the second matching result.

2. The method according to claim 1, characterized in that The target multimodal large language model is trained by the following steps: For each element that needs to be protected by copyright, obtain multiple pieces of training data containing the element and text description information of the element; The first associated data pair and the second associated data pair are associated and stored; wherein the first associated data pair is a data pair consisting of an element and text description information of the element; and the second associated data pair is a data pair consisting of training data and text description information of elements in the training data; The first associated data pair and / or the second associated data pair are input into an initial multimodal large language model for training to obtain the target multimodal large language model.

3. The method according to claim 2, characterized in that The copyright index library is obtained through the following steps: Inputting the elements in the second associated data pair into the target multimodal large language model for feature extraction to obtain element features; and inputting the text description information in the second associated data pair into the target multimodal large language model for feature extraction to obtain text description features; The element features and the text description features are associated and stored to form the copyright index library.

4. The method according to claim 1, characterized in that The elements that need to be protected by copyright include images, texts, audios, and videos.

5. The method according to claim 1, characterized in that The step of matching the target case feature with each piece of data in the created copyright index library to obtain a first matching result includes: Calculate the similarity between the target case feature and the element feature in each piece of data in the copyright index library; take the element features corresponding to the first N similarities in descending order as the first matching results; wherein N is a preset positive integer; and / or, The step of matching the target text description feature with each piece of data in the copyright index library to obtain a second matching result includes: Calculate the similarity between the target text description feature and the text description feature in each data in the copyright index library; and take the text description features corresponding to the first M similarities in descending order as the second matching results; wherein M is a preset positive integer.

6. The method according to claim 1, characterized in that The step of matching the target case feature with each piece of data in the created copyright index library to obtain a first matching result includes: Calculating the similarity between the target case feature and the element feature in each piece of data in the copyright index library; taking the element feature corresponding to the similarity greater than or equal to the first set threshold as the first matching result; and / or, The step of matching the target text description feature with each piece of data in the copyright index library to obtain a second matching result includes: Calculate the similarity between the target text description feature and the text description feature in each piece of data in the copyright index library; and take the element feature corresponding to the similarity greater than or equal to the second set threshold as the second matching result.

7. The method according to claim 5 or 6, characterized in that: The similarity is cosine similarity or Euclidean similarity.

8. A copyright identification device based on a multimodal large language model, characterized in that: The device comprises: An information acquisition module, used to acquire target case data to be subjected to copyright identification and / or target text description information of the target case data; the target case data is video data, audio data, image data or text data; A feature extraction module is used to input the target case data into a trained target multimodal large language model for feature extraction to obtain target case features of the target case data; and / or input the target text description information into the target multimodal large language model for feature extraction to obtain target text description features of the target text description information; wherein the target multimodal large language model is trained based on training data of multiple elements that need to be protected by copyright, and the element is the smallest unit that needs to be protected by copyright; A matching module is used to match the target case feature with each piece of data in the created copyright index library to obtain a first matching result; and / or, input the target text description feature into each piece of data in the copyright index library for matching to obtain a second matching result; the copyright index library includes: multiple pieces of data, each piece of data includes one-to-one corresponding element features and text description features; The infringement determination module is used to display the first matching result and / or the second matching result to determine whether the target case data is infringing based on the displayed first matching result and / or the second matching result.

9. An electronic device, characterized in that: include: a processor configured to execute program instructions; as well as A memory configured to store the program instructions, and when the program instructions are loaded and executed by the processor, the processor executes the copyright identification method based on a multimodal large language model according to any one of claims 1-7.

10. A computer-readable storage medium having program instructions stored therein, characterized in that: When the program instructions are loaded and executed by the processor, the processor executes the copyright identification method based on the multimodal large language model according to any one of claims 1 to 7.