Notification device, notification system, notification method, and computer program
The notification device uses a machine learning model to detect unauthorized content by identifying registered persons, prohibited phrases, and usage scopes, enhancing early detection and notification of unauthorized content for celebrity rights holders.
Patent Information
- Application Number
- JP2025121324
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2045-07-18
AI Technical Summary
It is difficult for rights holders of voices and images of celebrities to detect unauthorized content early among the vast amount of content available on the Internet due to advances in counterfeiting technologies like deepfakes.
A notification device equipped with an acquisition unit, first and second determination units, and a notification unit, utilizing a machine learning model to identify unauthorized content by determining if the content corresponds to registered persons, contains prohibited phrases, or exceeds permitted usage, and notifying unauthorized use.
Enables early detection and notification of unauthorized content, improving accuracy by considering digital watermarks, prohibited phrases, and permitted usage scopes, thereby protecting the rights of celebrities.
Smart Images

Figure 0007755106000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an alarm device. [Background technology]
[0002] In recent years, advances in counterfeiting technologies such as deepfakes have led to an increase in cases where fraudulent content is generated by processing or synthesizing data such as the voices and portraits of celebrities. As a countermeasure against counterfeiting technologies that generate such fraudulent content, for example, Patent Document 1 discloses a detection device that detects impersonation using deepfakes. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2024-002636 Summary of the Invention [Problem to be solved by the invention]
[0004] It is desirable for rights holders of the voices and images of celebrities (for example, celebrities themselves or their agencies) to detect inappropriate content early on. However, it is difficult to discover inappropriate content related to such rights holders among the vast amount of content available on the Internet.
[0005] One aspect of the present disclosure is to provide a technology that can detect unauthorized content related to a rights holder at an early stage. [Means for solving the problem]
[0006] One aspect of the present disclosure is a notification device including an acquisition unit, a first determination unit, a second determination unit, and a notification unit. The acquisition unit is configured to acquire content. The first determination unit is configured to determine whether a person associated with the content corresponds to at least one registered person registered in a database, based on an output of a machine learning model obtained by inputting the content into the trained machine learning model. The second determination unit is configured to determine whether the content is being misused. The notification unit is configured to notify that the content is misused content, on condition that it is determined that the person associated with the content corresponds to at least one registered person and that the content is being misused.
[0007] With this configuration, unauthorized content relating to the registered rights holder can be detected early.
[0008] In one aspect of the present disclosure, the second determination unit may be configured to determine that the content is being illegally used if at least one of a first condition, a second condition, and a third condition is satisfied. The first condition may be that the content does not have identification information associated with its authenticity. The second condition may be that the content contains a phrase included in the phrase list. The third condition may be that the scope of use of the content exceeds the permitted scope.
[0009] With this configuration, unauthorized content relating to the registered rights holder can be detected early.
[0010] In one aspect of the present disclosure, the second determination unit may be configured to determine that the content has been illegally used on condition that identification information relating to the authenticity of the content is not attached to the content.
[0011] According to this configuration, unauthorized content that does not have identification information attached thereto can be detected early.
[0012] The aspect of the present disclosure may further include a third determination unit configured to determine whether a source of the content is included in a predetermined list. The notification unit may be configured to notify that the content is unauthorized content based on the result of the determination by the third determination unit.
[0013] According to this configuration, whether or not the content is unauthorized can be determined based on whether or not the source of the content is included in a predetermined list, thereby enabling unauthorized content to be detected with greater accuracy.
[0014] In one aspect of the present disclosure, the notification unit may be configured not to notify that the content is unauthorized content when it is determined that the acquisition source is included in the list. With this configuration, even if the content does not have identification information attached, if the source of the content is included in a predetermined list, it is possible to prevent the content from being reported as unauthorized content as an exception.
[0015] In one aspect of the present disclosure, the list may include information indicating a platform set for one or more relevant persons as a source of acquisition, and the one or more relevant persons may be registered persons among at least one registered person who corresponds to a person associated with the content.
[0016] According to this configuration, it is possible to prevent content acquired from a platform corresponding to one or more relevant persons from being reported as unauthorized content. In one aspect of the present disclosure, the identification information may be a digital watermark.
[0017] According to this configuration, content that does not have a digital watermark attached can be reported as unauthorized content.
[0018] In one aspect of the present disclosure, the second determination unit may be configured to determine that the content is being misused on condition that the content contains a phrase included in the phrase list.
[0019] According to this configuration, content that includes a phrase included in the phrase list can be reported as unauthorized content, thereby enabling unauthorized content to be detected with greater accuracy.
[0020] In one aspect of the present disclosure, the list of phrases may include at least one of phrases contrary to public order and morals, phrases related to political activities, phrases related to religious activities, phrases related to the privacy of one or more relevant persons, and phrases set for each of one or more relevant persons. The one or more relevant persons may be registered persons who correspond to people associated with the content of at least one registered person.
[0021] According to this configuration, content that includes inappropriate language can be reported as unauthorized content, thereby enabling unauthorized content to be detected with greater accuracy.
[0022] In one aspect of the present disclosure, the second determination unit may be configured to determine that the content is being used illegally if the range of use of the content exceeds the permitted range. According to this configuration, unauthorized content that is being used beyond the permitted scope of the content can be detected early.
[0023] In one aspect of the present disclosure, the content may include audio and surrounding information, and the machine learning model may be configured to output information about a person associated with the content based on the audio and surrounding information. With this configuration, it is possible to detect unauthorized content that includes the voice of a registered user.
[0024] In one aspect of the present disclosure, the peripheral information may include information about the speaker of the voice. In one aspect of the present disclosure, the ambient information may include text associated with the speaker of the audio.
[0025] This configuration improves the accuracy of determining whether a person associated with content corresponds to at least one person registered in the database, thereby enabling more accurate detection of unauthorized content.
[0026] In one aspect of the present disclosure, the content may include a portrait and surrounding information, and the machine learning model may be configured to output information about a person associated with the content based on the portrait and surrounding information. With this configuration, it is possible to detect unauthorized content that includes a portrait of a registrant.
[0027] In one aspect of the present disclosure, the surrounding information may include information about the person corresponding to the portrait. In one aspect of the present disclosure, the surrounding information may include text associated with the person corresponding to the portrait.
[0028] This configuration improves the accuracy of determining whether a person associated with content corresponds to at least one person registered in the database, thereby enabling more accurate detection of unauthorized content.
[0029] One aspect of the present disclosure is a notification system including the above-described notification device, a database, and a display device. The display device is configured to be able to display information notified by the notification unit of the notification device.
[0030] With this configuration, it is possible to realize a system in which information notified by the notification device is displayed on the display device.
[0031] One aspect of the present disclosure is a notification method executed by a computer, which includes acquiring content, determining whether a person associated with the content corresponds to at least one registered person registered in a database based on the output of a machine learning model obtained by inputting the content into a trained machine learning model, determining whether the content is being misused, and notifying that the content is fraudulent content on the condition that it is determined that the person associated with the content corresponds to at least one registered person and that the content is being misused. This notification method provides the same effects as the notification device described above.
[0032] One aspect of the present disclosure is a computer program for causing a computer to execute processing, the processing including: acquiring content; determining whether a person associated with the content corresponds to at least one registered person registered in a database based on the output of a machine learning model obtained by inputting the content into a trained machine learning model; determining whether the content is being misused; and, if it is determined that the person associated with the content corresponds to at least one registered person and it is determined that the content is being misused, notifying that the content is misused content.
[0033] This computer program provides the same effects as the above-described notification device. [Brief explanation of the drawings]
[0034] [Figure 1] FIG. 2 is a block diagram showing the configuration of the notification device. [Figure 2] FIG. 2 is a diagram illustrating information held in a database. [Figure 3] 10 is a flowchart showing a notification process of the notification device. [Figure 4]10 is a flowchart showing a notification process for notifying unauthorized use of content as unauthorized content in another embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0035] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. [1. First embodiment] [1-1.Configuration] [1-1-1. Overall structure] The notification system 100 shown in Fig. 1 is a system for notifying content in which the voice, portrait, etc. of a celebrity is illegally used (hereinafter also referred to as illegal content). The notification system 100 is configured to notify the rights holder of the celebrity's voice, portrait, etc. of the illegal content. The rights holder may be, for example, the celebrity himself or the agency to which the celebrity belongs.
[0036] In this disclosure, "talent" refers to any person who appears in public and works through various media. Examples of entertainers include comedians, actors, voice actors, singers, idols, models, presenters, performers, announcers, athletes, fighters, shogi players, politicians, critics, experts, streamers (e.g., YouTubers), and influencers. Entertainers are not limited to real people, but may also be virtual characters. Examples of virtual characters include anime characters, game characters, mascot characters, virtual streamers (e.g., VTubers), and virtual idols.
[0037] Examples of media include mass media such as television and radio, and digital media that can distribute audio and video over the Internet. "Content" in this disclosure refers to digital content that includes at least one of audio, images, and video. Audio includes a person's voice. Images and video include a person's portrait. Content may also include text. Examples of text include the title of the content, a description, and comments posted by viewers of the content. Text may include words within an image or video.
[0038] The content may be various posts on various websites on the internet, such as social networking services (SNSs), audio platforms such as podcasts, image platforms such as image posting sites, and video platforms such as video posting sites.
[0039] The content may be text, audio, or the like generated for a given platform that allows two-way communication such as text chat or voice call.
[0040] In this embodiment, the content includes audio and peripheral information. The peripheral information is information other than the audio included in the content. For example, the peripheral information may include information about the speaker of the audio. The peripheral information may include text associated with the speaker of the audio. For example, the text associated with the speaker of the audio may include the speaker's name, occupation, titles of works in which the speaker has appeared, titles of programs in which the speaker has appeared, etc. The peripheral information may include metadata about the content.
[0041] The notification system 100 includes a notification device 1, a database 2, a machine learning model 3, and a display device 4. The notification device 1 is an information terminal such as a personal computer, a tablet terminal, etc. The notification device 1 includes a processor 11, a memory 12, a storage 13, and a communication interface .
[0042] The processor 11 is configured to execute processing in accordance with a computer program recorded in the storage 13 . The memory 12 is used as a work area when the processor 11 executes processing. Examples of the memory 12 include a random access memory (RAM), a read only memory (ROM), and a flash memory.
[0043] The storage 13 holds computer programs and data used when executing processes according to the computer programs. Examples of the storage 13 include a hard disk drive (HDD) and a solid state drive (SSD).
[0044] The communication interface 14 is an interface capable of communicating various types of data in accordance with a predetermined standard. The notification device 1 is configured to be able to communicate with each of the database 2, the machine learning model 3, and the display device 4 via the communication interface 14.
[0045] The notification device 1 is configured to acquire content through the communication interface 14, determine whether the acquired content is unauthorized content, and notify the unauthorized content by displaying the unauthorized content on the display device 4. Specific processing by the notification device 1 will be described later.
[0046] As shown in Fig. 2, Database 2 holds various types of information corresponding to at least one celebrity whose rights are protected. Rights holders can register various types of information corresponding to each celebrity in Database 2. In this embodiment, each registered celebrity is assigned a unique ID, and Database 2 is managed as a table in which various types of information corresponding to each celebrity are associated with each ID.
[0047] The database 2 has the following table items: ID 21, registrant 22, right holder 23, watermark information 24, prohibition message 25, permission range 26, permission target 27, and feature amount 28.
[0048] The ID21 is a unique value assigned to each registered talent. In the example of Figure 2, "0001", "0002", and "0003" are shown as the ID21. The registered users 22 indicate registered talents. In the example of Fig. 2, "Talent A", "Talent B", and "Talent C" are shown as the registered users 22.
[0049] The rights holder 23 indicates the rights holder corresponding to the registrant 22. In the example of Figure 2, "agency a," "talent B," and "agency c" are shown as the rights holders 23. For example, the rights holder 23 for "talent A" is "agency a."
[0050] The watermark information 24 is information for verifying whether the digital watermark attached to the content is valid. A digital watermark is identification information attached to content such as audio, images, video, and text. By using a digital watermark, it is possible to detect content tampering, protect copyrights, identify the source, etc. Hereinafter, a digital watermark will also be simply referred to as a watermark.
[0051] The watermark is attached to the content of the registrant 22 by the right holder 23. In other words, depending on whether or not the content has a proper watermark attached, it can be determined whether the content is legitimate content properly generated by the right holder 23 or unauthorized content illegally generated by someone other than the right holder. In the example of Figure 2, "watermark α," "watermark β," and "watermark γ" are shown as watermark information 24.
[0052] The prohibition wording 25 is a wording for determining whether the content is being used improperly. The prohibition wording 25 is used as a so-called blacklist. If the content contains the prohibition wording 25, the content is determined to be being used improperly (in other words, being used illegally). In the example of Figure 2, "wording X", "none", and "wording Y" are shown as the prohibition wording 25. "none" means that the prohibition wording 25 is not set.
[0053] The prohibited phrases 25 are, for example, phrases that are contrary to public order and morals, phrases related to political activities, phrases related to religious activities, phrases related to the privacy of celebrities, and phrases set for each celebrity.
[0054] Examples of language that violates public order and morals include language that contains sexual content, violent content, discriminatory content, illegal content, or anti-social content. Examples of statements related to political activities include statements supporting a particular political party, criticizing policies, and calling for demonstrations.
[0055] Examples of language related to religious activities include solicitations to join a particular religion, discrimination and slander against other religions or non-religious people, and language that attributes illness or disasters to divine punishment. Examples of language regarding a celebrity's privacy include descriptions of their body or health (such as their measurements or chronic illnesses), address, family composition, friendships, marital status, sexual orientation, gender identity, income, assets, debts, political or religious beliefs, etc.
[0056] Examples of the wording set for each talent include language that violates the talent's sponsorship contract (for example, language related to competitors), language that damages the talent's brand image (for example, abusive language or extreme language used by a talent with an image of purity), etc.
[0057] The prohibition message 25 may be set individually for each registrant 22, or may be set as a message common to a plurality of registrants 22.
[0058] The permitted range 26 is the range of content use permitted by the rights holder 23. The permitted range 26 is used to determine whether the range of content use is appropriate. The range of content use may include the form of the content. If the range of content use exceeds the permitted range 26, the content is determined to be used improperly (in other words, to be used illegally).
[0059] For example, if the rights holder 23 specifies that the scope of content use is "limited to use in audio content from Company P," then content used in products from Company Q, which is different from Company P, or content used in video content from Company P, etc., will be determined to be unauthorized use. In the example of Figure 2, "Company P," "audio," and "video" are shown as permitted scope 26. "Company P" indicates that the scope of content use is limited to use by Company P. "Audio" and "video" indicate that the scope of content use is limited to use in audio content and video content, respectively.
[0060] The permission target 27 is information indicating a proper source of content. The permission target 27 is used as a so-called whitelist. The permission target 27 may be information indicating a platform from which content is obtained. The permission target 27 is, for example, a URL, a domain name, an IP address, or the like.
[0061] The permission targets 27 are used to determine whether the source of the content is appropriate. If the source of the content is included in the permission targets 27, the source of the content is determined to be appropriate. In the example of FIG. 2, "none" and the domain name "example.com" are shown as the permission targets 27. "none" means that the permission targets 27 are not set.
[0062] The feature 28 is a feature for each talent that is used to allow the machine learning model 3 to infer the person associated with the content. In other words, the "person associated with the content" is the speaker in the content that is associated with the viewer of the content. The viewer of the content can associate who is speaking with the content based on the voice and portrait in the content. The "statement" is not limited to voice, but may also be expressed, for example, as text.
[0063] The feature 28 is, for example, an audio feature related to audio, an image feature related to an image, a video feature related to a video, a text feature related to text, a metadata feature related to metadata, and the like.
[0064] Examples of audio features include features related to the talent's voice, voiceprint, pitch, tone, speaking rate, speaking style, accent, and the like. Examples of image features and video features include features related to face, physique, clothing, movement, etc. Facial features include, for example, the positional relationship between the eyes, nose, and mouth, the proportions of the facial bone structure and contours, the color and shape of the eyes, and the pattern of facial expression.
[0065] Examples of text features include vocabulary, catchphrases, pronouns, word usage, topic trends, content titles, descriptions, and comments posted by viewers of the content. Examples of metadata features include the characteristics of the platform on which the content is posted, the tendency of the time period in which the content is posted, the poster of the content, and the like.
[0066] In the example of Figure 2, the features 28 are shown as "feature A1, feature A2, feature A3, ...", "feature B1, feature B2, feature B3, ...", and "feature C1, feature C2, feature C3, ...".
[0067] The machine learning model 3 is a trained machine learning model. The machine learning model 3 may be, for example, a deep learning model. The training method of the machine learning model 3 will be described later. The machine learning model 3 outputs information about a person associated with the input content. For example, the machine learning model 3 infers who the person associated with the content is based on information included in the input content, and outputs the inference result. In this embodiment, the machine learning model 3 outputs information about the person associated with the input content based on the voice and peripheral information included in the input content.
[0068] Furthermore, the machine learning model 3 may output the basis for the inference result. Examples of the basis for the inference result include the degree of coincidence of the features 28 (e.g., audio features), specific keywords included in the peripheral information, and words in the content that correspond to the prohibited words 25.
[0069] The person associated with the content in the inference result may be one person or two or more people. The inference result may be a combination of the person associated with the content and the probability that it is that person. For example, the inference result may be displayed as "Probability that it is person A is 70%, probability that it is person B is 24%."
[0070] The machine learning model 3 may infer who the content is associated with from among at least one registered person 22 registered in the database 2. The display device 4 is configured to be able to display information notified by the notification device 1. The display device 4 is, for example, a display.
[0071] As an example, the display device 4 outputs unauthorized content and information related to the unauthorized content. Examples of the information related to the unauthorized content include the source of the unauthorized content (e.g., URL), the acquisition date and time, the publication date and time, the inference result by the machine learning model 3 (e.g., the probability that it corresponds to the registrant 22), the basis for the inference result, etc. The rights holder 23 can detect unauthorized content through the information displayed on the display device 4.
[0072] [1-2. Processing] [1-2-1. Learning process] When a new talent is registered in the database 2 as a registrant 22, the processor 11 executes a learning process on the machine learning model 3. Hereinafter, a talent registered in the database 2 will also be referred to as a registered talent.
[0073] The processor 11 performs a learning process on the machine learning model 3 so as to estimate the probability that a person associated with a certain content corresponds to the registered talent, based on information contained in the content and feature quantities 28 related to the registered talent.
[0074] The machine learning model 3 extracts various features from the content and calculates the degree to which the person inferred from the various extracted features matches the person inferred from the features 28 related to the registered talent, thereby estimating the probability that the person associated with the content corresponds to the registered talent.
[0075] The processor 11 may perform supervised learning on the machine learning model 3 by preparing content as training data corresponding to each case so that the machine learning model 3 can make appropriate inferences about each of the following cases (A) to (C).
[0076] (A) Cases where the content explicitly misrepresents a registered talent Examples of such cases include when the name of the registered talent (full name, abbreviation, nickname, etc.) is clearly stated or declared in the content, or when the registered talent himself / herself is present in the image or video.
[0077] In this case, the correct label in the training data indicates that the person is a registered talent.
[0078] (B) Cases where the content implicitly misrepresents a registered talent This is the case, for example, when the content does not specify or state the name of the registered talent, but includes information related to the registered talent.
[0079] Examples of related information include information including occupations such as "veteran comedian," "famous singer," and "big-name actor," information including the names of films, dramas, anime, etc. in which the person has appeared, and information including the names of television, radio, online programs, etc. in which the person has appeared.
[0080] In this case, the correct label in the training data indicates that the person is a registered talent.
[0081] (C) Cases where the content explicitly denies that it is not a registered talent. This case includes, for example, cases where the content clearly states or declares that it is an imitation of a celebrity, or where the name of a person other than the registered celebrity is clearly stated or declared. Content of a person who simply has a similar voice to a registered celebrity falls into this category.
[0082] In this case, the correct label in the training data indicates "not a registered talent."
[0083] [1-2-2. Notification processing] The notification process executed by the processor 11 will be described with reference to the flowchart of FIG. The processor 11 starts the notification process shown in Fig. 3 at a predetermined cycle (for example, every hour). The processor 11 may start the notification process shown in Fig. 3 in response to a request from a user such as an administrator of the database 2 or a rights holder 23. The request may be received through a platform provided for notification. The request may include content.
[0084] First, in S100, the processor 11 acquires content. The processor 11 may acquire the content itself, or may acquire information indicating the location of the content, such as a URL. In this embodiment, the processor 11 acquires content that includes audio.
[0085] For example, the processor 11 uses a web crawler to acquire content including audio from a specific website. The web crawler may be a distributed crawler. The specific website may be, for example, various websites on the Internet, such as a social networking site (SNS), an audio platform, an image platform, or a video platform. The crawl target is not limited to a website, but may also be a mobile app such as a social networking site, an audio platform, an image platform, or a video platform.
[0086] Alternatively, if the notification process is initiated by a request from a user and the request includes content, the processor 11 may obtain the content from the request.
[0087] Next, in S110, the processor 11 performs preprocessing to process the acquired content so that it is easier for the machine learning model 3 to analyze. For example, the processor 11 performs noise removal, volume normalization, frame division, etc. on the audio included in the content. The processor 11 may extract the audio from the content. The processor 11 may extract text, images, videos, etc. from peripheral information about the content. The processor 11 may extract the audio from videos.
[0088] Next, in S120, the processor 11 uses the machine learning model 3 to determine whether the person associated with the content corresponds to at least one registrant 22 registered in the database 2. That is, the processor 11 inputs the content to the machine learning model 3 and receives, as an output from the machine learning model 3, an inference result that infers who the person associated with the content is. The processor 11 makes the determination in S120 based on the inference result of the machine learning model 3. The processor 11 may input the extracted audio and peripheral information to the machine learning model 3 in addition to or instead of the content. The processor 11 may input text, images, videos, etc. extracted from the peripheral information to the machine learning model 3 as the peripheral information.
[0089] If processor 11 determines in S120 that the person associated with the content does not correspond to any of registrants 22 registered in database 2 (S120: NO), processor 11 ends the notification process of FIG.
[0090] On the other hand, if the processor 11 determines in S120 that the person associated with the content corresponds to at least one registrant 22 registered in the database 2 (S120: YES), the process proceeds to S130.
[0091] In S130, the processor 11 determines whether or not the content has an appropriate watermark attached, based on the watermark information 24. The watermark information 24 used in S130 is the watermark information 24 corresponding to the registrant 22 determined to be applicable in S120.
[0092] If the processor 11 determines in S130 that the content has an appropriate watermark attached (S130: YES), the process proceeds to S140. In S140, the processor 11 determines whether the content contains any prohibited words 25. The prohibited words 25 used in S140 are the prohibited words 25 corresponding to the registrant 22 determined to be relevant in S120 and the prohibited words 25 common to multiple registrants 22. For example, when transcribing the audio included in the content, the processor 11 determines whether at least one prohibited word 25 is included in the transcribed text.
[0093] If the processor 11 determines in S140 that the content does not include the prohibition message 25 (S140: NO), the processor 11 proceeds to S150. In S150, the processor 11 determines whether the usage range of the content exceeds the content's permission range 26. The permission range 26 used in S150 is the permission range 26 corresponding to the registrant 22 determined to be applicable in S120.
[0094] If processor 11 determines in S150 that the usage range of the content does not exceed permitted range 26 of the content (S150: NO), processor 11 ends the notification process of FIG. On the other hand, if the processor 11 determines in S150 that the usage range of the content exceeds the permitted range 26 of the content (S150: YES), the process proceeds to S170.
[0095] In S170, the processor 11 notifies the user that the content is unauthorized content. That is, the processor 11 notifies the user of the unauthorized content and information related to the unauthorized content. The processor 11 may notify the user of the unauthorized content through the display device 4.
[0096] Thereafter, the processor 11 ends the notification process of FIG. On the other hand, if the processor 11 determines in S140 that the content contains the prohibition message 25 (S140: YES), the process proceeds to S170. The subsequent processing is the same as that described above.
[0097] On the other hand, if the processor 11 determines in S130 that the content does not have a proper watermark attached (S130: NO), the process proceeds to S160. In S160, the processor 11 determines whether the source of the content is included in the permission target 27. The permission target 27 used in S160 is the permission target 27 corresponding to the registrant 22 determined to be applicable in S120.
[0098] If the processor 11 determines in S160 that the content acquisition source is included in the permission targets 27 (S160: YES), the process proceeds to S140.
[0099] On the other hand, if the processor 11 determines in S160 that the content acquisition source is not included in the permission targets 27 (S160: NO), the process proceeds to S170. The subsequent processing is the same as that described above.
[0100] In this way, the processor 11 determines whether or not to report that the content is illegal content based on whether or not the person associated with the content corresponds to at least one registered person 22 registered in the database 2, whether or not the content is watermarked, whether or not the content contains prohibited wording 25, whether or not the scope of use of the content exceeds the permitted scope 26, and whether or not the source of the content is included in the permitted targets 27.
[0101] [1-3.Effects] According to the embodiment described above, the following actions and effects can be obtained. (1a) The processor 11 acquires content and, if it determines that the person associated with the acquired content corresponds to at least one registered person 22 registered in the database 2 and if it determines that the content does not have an appropriate watermark attached, reports that the content is unauthorized content.
[0102] This process allows the notification device 1 to extract content related to celebrities registered in the database 2 that does not have an appropriate watermark attached from the content acquired. Content that does not have an appropriate watermark attached is likely to be unauthorized content. This allows the rights holder 23 to detect unauthorized content related to celebrities registered in the database 2 at an early stage.
[0103] (1b) The rights holder 23 can select a countermeasure for the detected unauthorized content. For example, if the rights holder 23 checks the detected unauthorized content and determines that it is problematic, it can request the source, creator, provider, etc. of the content to delete or modify the content.
[0104] When making the request, the rights holder 23 may include in the request the basis for the inference result by the machine learning model 3. The rights holder 23 may obtain the basis for the inference result from information notified by the notification device 1.
[0105] This process can improve the validity of requests to delete or modify content.
[0106] Alternatively, if the rights holder 23 checks the detected unauthorized content and determines that there is no problem, the rights holder 23 may choose to take no action as a countermeasure. According to this processing, even if the notification device 1 mistakenly reports content that is not illegal as illegal content, if the rights holder 23 determines that there is no problem, the content can be prevented from being mistakenly deleted.
[0107] (1c) If it is determined that the person associated with the content does not correspond to any of the registered persons 22 registered in the database 2, the content is not reported as unauthorized content.
[0108] According to this process, content relating to a celebrity who is not registered in the database 2 is not reported as unauthorized content, which makes it possible to prevent the rights holder 23 from being excessively notified.
[0109] (1d) The content input to the machine learning model 3 includes audio. Based on the audio included in the input content, the machine learning model 3 outputs information about a person associated with the content.
[0110] This process makes it possible to notify unauthorized content that includes the voice of a registered celebrity, thereby enabling unauthorized content related to the registered celebrity to be detected early.
[0111] (1e) The content input to the machine learning model 3 includes audio and peripheral information. The machine learning model 3 outputs information about the person associated with the input content based on the audio and peripheral information contained in the input content. For example, the machine learning model 3 predicts the person associated with the content based not only on the voice of the person speaking in the content, but also on peripheral information such as the title of the content, description, and comments posted by viewers of the content.
[0112] According to this process, the machine learning model 3 can make inferences using not only the voice but also the peripheral information, thereby improving the inference accuracy of the machine learning model 3. As a result, it is possible to improve the accuracy of determining whether a person associated with content corresponds to at least one registrant 22 registered in the database 2.
[0113] (1f) The processor 11 determines whether or not to report the content as unauthorized content based on whether or not the content has a proper watermark attached thereto. According to such a process, for example, by attaching an appropriate watermark to legitimate content generated by the rights holder 23, it is possible to easily determine whether the content is legitimate (in other words, whether the content is unauthorized).
[0114] (1g) The processor 11 notifies the user that the content is fraudulent content if it is determined that the person associated with the content corresponds to at least one registered user 22 registered in the database 2 and that the content contains prohibited wording 25.
[0115] According to this process, even if a content has a proper watermark, if the content contains the prohibition message 25, the content can be determined to be illegally used and reported as illegal content. This allows the rights holder 23 to detect illegal content containing the prohibition message 25 at an early stage.
[0116] (1h) The processor 11 notifies the user that the content is unauthorized content if it is determined that the person associated with the content corresponds to at least one registered user 22 registered in the database 2 and that the scope of use of the content exceeds the permitted scope 26 of the content.
[0117] According to this process, even if the content has a proper watermark attached, if the range of use of the content exceeds the content's permitted range 26, the content can be determined to be being used illegally and reported as unauthorized content. This allows the rights holder 23 to quickly detect unauthorized content that is being used beyond the content's permitted range 26.
[0118] (1i) If the content does not have a proper watermark attached, the processor 11 notifies the user that the content is unauthorized content, provided that the source of the content is not included in the permitted targets 27.
[0119] According to this process, even if the content does not have a proper watermark, the processor 11 exceptionally does not report that the content is unauthorized content if the source of the content is included in the permission target 27. This allows the rights holder 23 to prevent content placed in the permission target 27 from being reported as unauthorized content.
[0120] [1-4. Correspondence between terms] In the above embodiment, the prohibited phrase 25 corresponds to an example of a phrase list, and the permitted target 27 corresponds to an example of a predetermined list.
[0121] The processing of S100 corresponds to an example of processing executed as an acquisition unit, the processing of S120 corresponds to an example of processing executed as a first judgment unit, the processing of S130, S140, and S150 corresponds to an example of processing executed as a second judgment unit, the processing of S160 corresponds to an example of processing executed as a third judgment unit, and the processing of S170 corresponds to an example of processing executed as an alarm unit.
[0122] 2. Other Embodiments Although the embodiments of the present disclosure have been described above, it goes without saying that the present disclosure is not limited to the above-described embodiments and can take on various forms.
[0123] (2a) In the above embodiment, the processor 11 executes the processes from S130 to S160 after executing the process of S120. However, the timing at which the processor 11 executes the process of S120 is not limited to before the process of S130. For example, the processor 11 may execute the process of S120 after S130 and before S170. In other words, the processor 11 may determine whether or not the content has an appropriate watermark, and then determine whether or not the person associated with the content corresponds to at least one registrant 22 registered in the database 2.
[0124] (2b) In the above embodiment, if the processor 11 determines in S160 that the content acquisition source is included in the permission targets 27 (S160: YES), the process proceeds to S140. However, if the processor 11 determines in S160 that the content acquisition source is included in the permitted targets 27 (S160: YES), the processor 11 may proceed to S150 or may end the notification process shown in FIG.
[0125] (2c) In the above embodiment, in S110, processor 11 performs preprocessing to process the acquired content so that it is easier for machine learning model 3 to analyze. However, processor 11 does not have to perform the process of S110. In other words, processor 11 may input content that has not been preprocessed to machine learning model 3.
[0126] (2d) In the above embodiment, the processor 11 determines the authenticity of the content in S130 by determining whether the content has a proper watermark attached thereto. However, the method for determining the authenticity of the content is not limited to the method using a watermark. For example, the processor 11 may determine the authenticity of the content by using a technology such as a digital signature or a blockchain technology.
[0127] (2e) In the above embodiment, the machine learning model 3 infers one or more persons who correspond to the person associated with the content from among at least one registrant 22 registered in the database 2. That is, in the notification system 100 of the above embodiment, one machine learning model 3 infers one or more persons from among at least one registrant 22. However, the number of machine learning models 3 included in the notification system 100 is not limited to one. For example, the notification system 100 may include, for each of at least one registrant 22, a machine learning model 3 that infers whether the person associated with the content corresponds to one registrant 22.
[0128] (2f) In the above embodiment, the content includes audio and peripheral information. The peripheral information is information other than audio included in the content. However, the content is not limited to content that includes audio.
[0129] For example, the content may include an image and peripheral information. In this case, the peripheral information may be information other than the image included in the content. The image may include a portrait. The peripheral information may be information other than the portrait included in the content. The peripheral information may include information about a person corresponding to the portrait. The peripheral information may include text associated with the person corresponding to the portrait. The person associated with the content may be a person associated with the portrait.
[0130] This process makes it possible to notify users of content that includes portraits of registered celebrities, thereby enabling early detection of unauthorized content related to registered celebrities.
[0131] (2g) The right holder 23 who has been notified of the infringing content can choose how to deal with the infringing content. For example, the right holder 23 can choose to request that the content be deleted or modified, or to do nothing.
[0132] In response to this, the notification system 100 may be configured to acquire the countermeasure selection result and provide feedback to the machine learning model 3. For example, if a request to delete or modify content is made as a countermeasure, the notification system 100 may provide feedback to the machine learning model 3 that the prediction result of the machine learning model 3 was valid. Alternatively, if no countermeasure is taken, the notification system 100 may provide feedback to the machine learning model 3 that the prediction result of the machine learning model 3 was invalid. By performing such processing, the estimation accuracy of the machine learning model 3 can be improved.
[0133] (2h) In the above embodiment, the processor 11 determines whether to proceed to S170 (in other words, whether to notify that the content is unauthorized content) based on the determination results of S130, S140, S150, and S160.
[0134] However, the processor 11 may determine whether to proceed to S170 based on the result of determining whether the content is being used illegally. That is, as shown in Fig. 4, the processor 11 may execute the process of S230 instead of S130, S140, S150, and S160.
[0135] In S230, the processor 11 may determine whether the content is being illegally used. If the processor 11 determines that the content is being illegally used (S230: YES), the processor 11 may proceed to S170 and notify the user that the content is illegal content.
[0136] On the other hand, if processor 11 determines that the content has not been used illegally (S230: NO), it may end the processing shown in FIG. The method for determining whether content is being used illegally is not particularly limited. For example, the processor 11 may execute at least one of the processes of S130, S140, and S150 as a process for determining whether content is being used illegally. The order of the processes of S130, S140, and S150 is not particularly limited, and may be executed in any order.
[0137] Alternatively, the processor 11 may determine whether or not the content is being misused by a method other than S130, S140, and S150. For example, the processor 11 may determine whether or not the content is being misused by using the machine learning model 3. As an example, the processor 11 may instruct the machine learning model 3 to infer whether or not the content includes inappropriate language, and determine whether or not the content is being misused based on the output from the machine learning model 3.
[0138] The processor 11 may further execute the process of S160 in addition to at least one of the processes of S130, S140, and S150. According to this process, processor 11 acquires content and, if it is determined that the person associated with the acquired content corresponds to at least one registrant 22 registered in database 2 and that the content is being used illegally, notifies the right holder 23 that the content is illegal. This allows the right holder 23 to quickly detect illegal content related to a talent registered in database 2.
[0139] (2i) In the above embodiment, the learning of the machine learning model 3 is performed by the learning process of the processor 11. However, the learning of the machine learning model 3 is not limited to this. For example, the notification system 100 may be provided with an existing trained machine learning model, such as a general-purpose large-scale language model, as the machine learning model 3.
[0140] (2j) Multiple functions possessed by one component in the above embodiments may be realized by multiple components, or one function possessed by one component may be realized by multiple components. Multiple functions possessed by multiple components may be realized by one component, or one function realized by multiple components may be realized by one component. Part of the configuration of the above embodiments may be omitted. At least part of the configuration of the above embodiments may be added to or substituted for the configuration of another of the above embodiments.
[0141] (2k) The present disclosure may be realized in various forms in addition to the above-described notification device, such as a system including the notification device as a component, a computer program for causing a computer to function as the notification device, a non-transitory tangible recording medium such as a semiconductor memory on which the computer program is recorded, a notification method, etc.
[0142] [Technical idea disclosed in this specification] [Item 1] An alarm device, an acquisition unit configured to acquire content; a first determination unit configured to determine whether a person associated with the content corresponds to at least one registered person registered in a database, based on an output of the machine learning model obtained by inputting the content into the trained machine learning model; a second determination unit configured to determine whether the content is being misused; a notification unit configured to notify that the content is unauthorized content on condition that it is determined that the person corresponds to the at least one registered user and that the content is being illegally used; An alarm device comprising:
[0143] [Item 2] The alarm device according to item 1, the second determination unit is configured to determine that the content is being illegally used if at least one of a first condition, a second condition, and a third condition is satisfied; The first condition is that the content does not have any identification information regarding the authenticity of the content attached thereto; The second condition is that the content contains a phrase included in the phrase list. The third condition is that the scope of use of the content exceeds the scope of the license. Alarm device.
[0144] [Item 3] The notification device according to item 1 or 2, the second determination unit is configured to determine that the content has been illegally used on condition that identification information regarding the authenticity of the content is not attached to the content. Alarm device.
[0145] [Item 4] The notification device according to item 3, a third determination unit configured to determine whether a source of the content is included in a predetermined list; the notification unit is configured to notify that the content is unauthorized content further based on the result of the determination made by the third determination unit. Alarm device.
[0146] [Item 5] Item 4. The alarm device according to item 4, the notification unit is configured not to notify that the content is unauthorized content when it is determined that the acquisition source is included in the list. Alarm device.
[0147] [Item 6] The notification device according to item 4 or 5, the list includes information indicating a platform set for one or more relevant persons as the source of the acquisition; The one or more relevant persons are registrants corresponding to the person among the at least one registrant; Alarm device.
[0148] [Item 7] The notification device according to any one of items 2 to 6, the identification information is a digital watermark; Alarm device.
[0149] [Item 8] The notification device according to any one of items 1 to 7, the second determination unit is configured to determine that the content is being misused on condition that a phrase included in the phrase list is included in the content. Alarm device.
[0150] [Item 9] The notification device according to item 2 or 8, The list of phrases includes at least one of a phrase against public order and morals, a phrase related to political activities, a phrase related to religious activities, a phrase related to the privacy of one or more persons, and a phrase set for each of the one or more persons; The one or more relevant persons are registrants corresponding to the person among the at least one registrant; Alarm device.
[0151] [Item 10] The notification device according to any one of items 1 to 9, the second determination unit is configured to determine that the content is being used illegally on condition that the scope of use of the content exceeds a permitted scope. Alarm device.
[0152] [Item 11] The notification device according to any one of items 1 to 10, the content includes audio and peripheral information; The machine learning model is configured to output information about the person based on the audio and the surrounding information. Alarm device.
[0153] [Item 12] Item 11: The alarm device according to item 11, The peripheral information includes information about a speaker of the voice. Alarm device.
[0154] [Item 13] The notification device according to item 11 or 12, the peripheral information includes text associated with the speaker of the audio; Alarm device.
[0155] [Item 14] The notification device according to any one of items 1 to 10, the content includes a portrait and surrounding information; the machine learning model is configured to output information about the person based on the portrait and the surrounding information. Alarm device.
[0156] [Item 15] Item 14: The alarm device according to item 14, The surrounding information includes information about the person corresponding to the portrait. Alarm device.
[0157] [Item 16] Item 14 or 15. The notification device according to item 14 or 15, the surrounding information includes text associated with the person corresponding to the portrait; Alarm device.
[0158] [Item 17] An alarm device according to any one of items 1 to 16; the database; a display device configured to display information notified by the notification unit of the notification device; An alarm system comprising:
[0159] [Item 18] A computer-implemented notification method, comprising: Obtaining content; inputting the content into a trained machine learning model and determining, based on an output of the machine learning model, whether or not a person associated with the content corresponds to at least one registered person registered in a database; determining whether the content is being misused; notifying the user that the content is unauthorized content on condition that it is determined that the person corresponds to the at least one registered user and that the content is being unauthorizedly used; A notification method including:
[0160] [Item 19] A computer program for causing a computer to execute a process, The process comprises: Obtaining content; inputting the content into a trained machine learning model and determining, based on an output of the machine learning model, whether or not a person associated with the content corresponds to at least one registered person registered in a database; determining whether the content is being misused; notifying the user that the content is unauthorized content on condition that it is determined that the person corresponds to the at least one registered user and that the content is being unauthorizedly used; a computer program comprising: [Explanation of symbols]
[0161] 1...alarm device, 2...database, 3...machine learning model, 4...display device, 11...processor, 100...alarm system
Claims
1. An alarm device, an acquisition unit configured to acquire content; a first determination unit configured to determine whether a person associated with the content corresponds to at least one registered person registered in a database, based on an output of the machine learning model obtained by inputting the content into the trained machine learning model; a second determination unit configured to determine whether the content is being used fraudulently; a notification unit configured to notify that the content is unauthorized content on condition that it is determined that the person corresponds to the at least one registered user and that the content is being unauthorizedly used; An alarm device comprising:
2. The notification device according to claim 1, the second determination unit is configured to determine that the content is being illegally used when at least one of a first condition, a second condition, and a third condition is satisfied; The first condition is that the content does not have any identification information regarding the authenticity of the content attached thereto; The second condition is that the content contains a phrase included in the phrase list. The third condition is that the scope of use of the content exceeds the permitted scope. Alarm device.
3. The notification device according to claim 1, the second determination unit is configured to determine that the content has been illegally used on condition that identification information regarding the authenticity of the content is not attached to the content. Alarm device.
4. The notification device according to claim 3, a third determination unit configured to determine whether a source of the content is included in a predetermined list; the notification unit is configured to notify that the content is unauthorized content further based on the result of the determination made by the third determination unit. Alarm device.
5. The notification device according to claim 4, the notification unit is configured not to notify that the content is unauthorized content when it is determined that the acquisition source is included in the list. Alarm device.
6. The notification device according to claim 5, the list includes information indicating a platform set for one or more relevant persons as the source of the acquisition; The one or more relevant persons are registrants corresponding to the person among the at least one registrant; Alarm device.
7. The notification device according to claim 3, the identification information is a digital watermark; Alarm device.
8. The notification device according to claim 1, the second determination unit is configured to determine that the content is being misused on condition that a phrase included in the phrase list is included in the content; Alarm device.
9. The notification device according to claim 8, the list of phrases includes at least one of a phrase contrary to public order and morals, a phrase related to political activities, a phrase related to religious activities, a phrase related to the privacy of one or more persons, and a phrase set for each of the one or more persons; The one or more relevant persons are registrants corresponding to the person among the at least one registrant; Alarm device.
10. The notification device according to claim 1, the second determination unit is configured to determine that the content is being used illegally on condition that the scope of use of the content exceeds a permitted scope. Alarm device.
11. The notification device according to claim 1, the content includes audio and peripheral information; The machine learning model is configured to output information about the person based on the audio and the surrounding information. Alarm device.
12. The notification device according to claim 11, The peripheral information includes information about a speaker of the voice. Alarm device.
13. The notification device according to claim 11, the peripheral information includes text associated with the speaker of the audio; Alarm device.
14. The notification device according to claim 1, the content includes a portrait and surrounding information; the machine learning model is configured to output information about the person based on the portrait and the surrounding information. Alarm device.
15. The notification device according to claim 14, The surrounding information includes information about the person corresponding to the portrait. Alarm device.
16. The notification device according to claim 14, the surrounding information includes text associated with the person corresponding to the portrait; Alarm device.
17. The notification device according to any one of claims 1 to 16, the database; a display device configured to display information notified by the notification unit of the notification device; An alarm system comprising:
18. A computer-implemented notification method, comprising: Obtaining content; inputting the content into a trained machine learning model and determining, based on an output of the machine learning model, whether or not a person associated with the content corresponds to at least one registered person registered in a database; determining whether the content is being misused; notifying the user that the content is unauthorized content on condition that it is determined that the person corresponds to the at least one registered user and that the content is being unauthorizedly used; A notification method including:
19. A computer program for causing a computer to execute a process, The process comprises: Obtaining content; inputting the content into a trained machine learning model and determining, based on an output of the machine learning model, whether or not a person associated with the content corresponds to at least one registered person registered in a database; determining whether the content is being misused; notifying the user that the content is unauthorized content on condition that it is determined that the person corresponds to the at least one registered user and that the content is being unauthorizedly used; a computer program comprising:
Citation Information
Patent Citations
Image content providing system, providing server, and image content providing method
JP2012078884A
Pretending detection program, detector, and method for detecting pretending
JP2024002636A
Avatar management system, avatar management method, program, and computer-readable recording medium
JP7485235B2
Methods for identifying, disrupting and monetizing the illegal sharing and viewing of digital and analog streaming content
US10080047B1
JPP7485235B