Content monitoring device and program
The content monitoring system addresses inefficiencies in manual keyword submission by automatically estimating and monitoring content legitimacy through HTML analysis and learning models, enhancing detection of unauthorized content.
Patent Information
- Application Number
- JP2022024906
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2042-02-21
AI Technical Summary
Existing content monitoring systems require manual submission of keywords for new content, which is cumbersome and inefficient.
A content monitoring device and program that automatically estimates legitimate content by analyzing web page data, including HTML tags and images, and using learning models to identify content, and determines legitimacy of content on other websites.
Automatically estimates and monitors content legitimacy, reducing manual input and efficiently detecting unauthorized content across multiple websites.
Smart Images

Figure 0007794016000001 
Figure 0007794016000002 
Figure 0007794016000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a content monitoring device and a program. [Background technology]
[0002] With the widespread use of the Internet, various digital content (hereinafter simply referred to as "content") such as digital comics (digital manga), animations, movies, and other videos can now be viewed by downloading or streaming from websites. On the other hand, there are sites that upload illegal content, such as so-called pirate sites. When illegal content is uploaded to pirate sites, publishers, distributors, copyright holders, etc. who publish the content are unable to obtain profits (such as copyright royalties) that they would otherwise receive, which has become a social problem.
[0003] In view of this situation, for example, a patent document describes a "content illegal distribution countermeasure system comprising: a monitoring request receiving means for receiving requests from customers regarding content to be monitored; a content monitoring means for monitoring the posting of content, including content that has been altered as illegal content, on an Internet content sharing site based on keywords; a posting status notifying means for notifying the customer of the posting status when the content monitoring means detects at least the posting of the illegal content; a pseudo-illegal content generating means for generating pseudo-illegal content that is similar in appearance to the posted illegal content, such as posting notation, but whose content is different from the illegal content; and a content distributing means for distributing the pseudo-illegal content to an Internet content sharing site" (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 5413028 Publication Summary of the Invention [Problem to be solved by the invention]
[0005] It is possible to search for illegal content using the method described in Patent Document 1, etc. When searching for illegal content, it is common to search for illegal content based on keywords or the like that indicate legitimate content submitted by the customer, as in the method described in Patent Document 1. Therefore, whenever a customer uploads new content, the customer must submit keywords or the like for the new content each time, which is cumbersome.
[0006] Therefore, an object of the present invention is to provide a content monitoring device and program that automatically estimates legitimate content and utilizes the estimation for monitoring content. [Means for solving the problem]
[0007] The present invention solves the above problems by the following means. A first invention is a content monitoring device comprising: a base value receiving means for receiving a base value including an input value that identifies content and address information of a web page on which the content is posted; and a content estimation means for estimating content posted on a web page from the web page corresponding to the address information of the base value received by the base value receiving means, and acquiring information related to the content of the estimated content. A second invention is a content monitoring device according to the first invention, further comprising a collection means for crawling websites where content can be viewed and / or downloaded and collecting other web pages on which the content estimated by the content estimation means is posted, and a content determination means for determining whether the content of the other websites collected by the collection means is legitimate content based on information relating to the content acquired by the content estimation means. A third invention is a content monitoring device according to the second invention, further comprising a notification means for notifying the address information of the other website collected by the collection means when the content determination means determines that the content is not legitimate. A fourth invention is a content monitoring device in any one of the first to third inventions, wherein the content estimation means refers to a markup description of the web page corresponding to the address information of the basic value accepted by the basic value accepting means, estimates the content posted on the web page from tag information related to the input value, and acquires information related to the content. A fifth invention is a content monitoring device in any one of the first to third inventions, wherein the content estimation means performs character recognition on a content image of the web page corresponding to the address information of the basic value accepted by the basic value accepting means, estimates the content posted on the web page, and acquires information related to the content. A sixth invention is a content monitoring device according to any one of the first to third inventions, wherein the content estimation means inputs the web page corresponding to the address information of the base value accepted by the base value accepting means and the input value into a learning device that has trained using a web page on which content is posted and an input value that identifies the content, thereby estimating the content posted on the input web page and obtaining information related to the content. A seventh invention is a content monitoring device according to the sixth invention, wherein the learning device is a model trained based on features using images of web pages on which the content is posted. An eighth invention is a content monitoring device according to the sixth invention, wherein the learning device is a model trained using connections between character strings and tag information indicated in markup descriptions of web pages on which the content is posted. A ninth invention is a content monitoring device according to any one of the first to eighth inventions, further comprising: a page registration means for storing the web page indicated by the address information of the basic value received by the basic value receiving means in a page memory unit; an update detection means for periodically checking the web page corresponding to the address information of the basic value received by the basic value receiving means and detecting updates to the web page stored in the page memory unit; and a page update means for updating the web page stored in the page memory unit to the updated web page detected by the update detection means, wherein the content estimation means estimates content published on the updated web page from the updated web page detected by the update detection means and acquires information related to the content. A tenth aspect of the present invention is a program for causing a computer to function as any one of the content monitoring devices according to the first to ninth aspects of the present invention. [Effects of the Invention]
[0008] According to the present invention, it is possible to provide a content monitoring device and a program that automatically estimate legitimate content and utilize the estimation for monitoring content. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is an overall configuration diagram of a content monitoring system according to an embodiment of the present invention and a functional block diagram of a content monitoring server. [Figure 2] 10 is a flowchart showing a main process in the content monitoring server according to the embodiment. [Figure 3] 10 is a flowchart showing a content estimation process in the content monitoring server according to the embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of HTML data of a web page according to the present embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of image data of a web page according to the present embodiment. [Figure 6]10 is a flowchart showing a page update confirmation process in the content monitoring server according to the embodiment. [Figure 7] 10 is a flowchart showing a content monitoring process in the content monitoring server according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, this is merely an example, and the technical scope of the present invention is not limited to this example. (Embodiment) <Content Monitoring System 100> FIG. 1 is a diagram showing the overall configuration of a content monitoring system 100 according to this embodiment and a functional block diagram of a content monitoring server 1. The content monitoring system 100 is a system that, in response to a request from a customer, estimates the content published on a web page designated by the customer, such as the customer's homepage, quickly finds illegal content on other websites, and provides the information to the customer.
[0011] The content monitoring system 100 includes a content monitoring server 1 (content monitoring device), a customer terminal 4, and a web server 7. The content monitoring server 1, the customer terminal 4, and the web server 7 are each connected to each other via a communication network N so that they can communicate with each other.
[0012] <Customer terminal 4> The customer terminal 4 is, for example, a terminal provided at a publisher, a distribution company, etc. In the following description, a company such as a publisher or a distribution company that requests content monitoring by the content monitoring system 100 will be referred to as a customer. A customer requests the content monitoring server 1 to monitor whether any of the content posted on their own homepage or the like has been illegally uploaded to other websites or the like. The customer uses, for example, the customer terminal 4 to send request information (base value) to the content monitoring server 1, which includes the URL (address information) of the web page on which the customer's content is posted and the title name (input value) of at least one piece of content posted on the web page. The customer need only make this request, including the web page URL, once at the beginning for each web page URL. Furthermore, the customer terminal 4 receives from the content monitoring server 1 the URL on which the content determined to be unauthorized content is posted.
[0013] The customer terminal 4 is, for example, a personal computer (PC), a smartphone, etc. In Fig. 1, the customer terminal 4 is illustrated as a notebook PC. Although not shown, the customer terminal 4 includes a control unit, a storage unit, an input unit, a display unit, a communication interface unit, etc. The customer terminal 4 may also include a touch panel display in which the input unit and the display unit are integrated.
[0014] <Web Server 7> The web server 7 is a server that provides websites to users, such as, for example, websites from all over the world, particularly websites where content can be viewed and / or downloaded. Although not shown, the web server 7 includes a control unit, a storage unit, a communication interface unit, and the like.
[0015] <Content monitoring server 1> The content monitoring server 1 estimates the content posted on the web page based on the URL of the web page requested by the customer and the title of the content. The content monitoring server 1 also crawls web pages on the website to collect other web pages that post the estimated content. The content monitoring server 1 then determines whether the content posted on the collected other web pages is legitimate or not, and if it determines that the content is not legitimate, notifies the customer terminal 4 of the URL of the other website. The content monitoring server 1 is, for example, a server operated by a consulting company. The content monitoring server 1 may be configured, for example, by one server, by a plurality of servers, or by a cloud.
[0016] As shown in FIG. 2, the content monitoring server 1 includes a control unit 10, a storage unit 20, and a communication interface unit 29. The control unit 10 is a CPU (Central Processing Unit) that controls the entire content monitoring server 1. The control unit 10 appropriately reads and executes an OS (Operating System) and application programs stored in the storage unit 20, thereby cooperating with the above-mentioned hardware and executing various functions.
[0017] The control unit 10 includes a request information receiving unit 11 (basic value receiving means), a page registration unit 12 (page registration means), a content estimation unit 13 (content estimation means), a crawling unit 14 (collection means), a judgment unit 15 (content judgment means), a notification unit 16 (notification means), an update detection unit 17 (update detection means), and a page update unit 18 (page update means).
[0018] The request information receiving unit 11 receives request information including the title of the content and the URL of the web page on which the customer's content is posted by receiving it from the customer terminal 4. Here, the title of the content is an input value that identifies the content. The web page URL is, for example, the address information of the customer's homepage. The web page URL may also be the address information of an official web page on which the customer has requested the content to be posted.
[0019] Page registration unit 12 stores the web page corresponding to the URL included in the request information in page storage unit 23, which will be described later. Page registration unit 12 may, for example, register image data of the web page (see FIG. 5, which will be described later), or may register HTML (markup description) data, which is the source code of the web page (see FIG. 4, which will be described later).
[0020] The content estimation unit 13 estimates the content posted on the web page from the web page corresponding to the URL included in the request information. The content estimation unit 13, for example, refers to the HTML data of the web page corresponding to the URL and estimates the content from tag information related to the title of the content. Furthermore, the content estimation unit 13 may, for example, perform character recognition on the content image of the web page corresponding to the URL and estimate the content from the relationship between the title of the content and the content image. Furthermore, the content estimation unit 13 may use a learning model (learning device) to estimate the content posted on the web page corresponding to the URL.
[0021] Then, the content estimation unit 13 acquires information related to the estimated content. The information related to the content may be, for example, the title name, publisher name, author name, etc. of the content, or tag information including the file name and location information of the content. Furthermore, when detecting an update to the web page corresponding to the URL included in the request information, the content estimation unit 13 estimates the content posted on the web page from the updated web page. At this time, the content estimation unit 13 uses the estimation method used when estimating the content posted on the web page from the web page corresponding to the URL included in the request information.
[0022] The crawling unit 14 crawls the website provided by the web server 7. Then, the crawling unit 14 collects other web pages on which the content estimated by the content estimation unit 13 is posted. Here, the website crawled by the crawling unit 14 is, as explained above with respect to the web server 7, for example, a website where content can be viewed and / or downloaded. The content monitoring system 100 may limit the websites crawled by the crawling unit 14 to specific websites, such as "pirated sites" or sites dedicated to videos, such as YouTube (registered trademark).
[0023] The determination unit 15 determines whether the content of other websites collected by the crawling unit 14 is legitimate content based on the information related to the content acquired by the content estimation unit 13. Here, content managed by a customer is called legitimate content. On the other hand, content not managed by a customer, such as on a pirated site, is not legitimate content. Such content is called illegal content, which is different from legitimate content. The determination unit 15 may identify, among other content determined to be unauthorized content, content whose title matches or is similar to that of the authorized content as illegal content to be monitored. Furthermore, the determination unit 15 may uniformly identify, for example, all content, including other content, on a website containing other content determined to be unauthorized content as illegal content. The determination unit 15 may then manageably store information related to the identified illegal content in, for example, the storage unit 20.
[0024] The notification unit 16 notifies the customer terminal 4 of the URL where the other content that the determination unit 15 has determined not to be legitimate content is posted. Here, the notification unit 16 may, for example, report the web page where the illegal content is posted to an organization that officially manages illegal content, such as a national branch office (management center) that regulates illegal content. Note that the notification method may be, for example, by email.
[0025] Update detection unit 17 periodically checks the web page corresponding to the URL included in the request information received by request information receiving unit 11. Update detection unit 17 then compares the web page with the web page stored in page storage unit 23 to detect updates to the web page stored in page storage unit 23 (described later). The page update unit 18 updates the web page stored in the page storage unit 23 to the detected updated web page.
[0026] The storage unit 20 is a storage area such as a hard disk or semiconductor memory element for storing programs, data, etc. required for the control unit 10 to execute various processes. The storage unit 20 includes a program storage unit 21, a content information storage unit 22, and a page storage unit 23. The program storage unit 21 is a storage area for storing programs. The program storage unit 21 stores various programs, including programs for executing various functions of the control unit 10. Note that the program for executing various functions of the control unit 10 may be one program, or different programs may be used for each functional unit, or multiple functional units may form one program.
[0027] The content information storage unit 22 is a storage area that stores information related to content, including request information obtained from a customer. The content information storage unit 22 stores request information including, for example, the URL of a web page and the title of the content. The content information storage unit 22 also stores information related to content estimated by the content estimation unit 13. Furthermore, the content information storage unit 22 may store information related to the estimation method used by the content estimation unit 13 when making the estimation. The content information storage unit 22 stores items included in the request information, such as a URL and a title name of the content, and information related to the content, using, for example, a request ID (IDentification) that identifies the request information as a key.
[0028] The page storage unit 23 is a storage area for storing web pages corresponding to URLs included in request information obtained from customers. The page storage unit 23 stores image data and HTML data of web pages, for example, using a request ID as a key.
[0029] The communication interface unit 29 is an interface for communicating with the customer terminal 4, the web server 7, etc. via the communication network N. Here, a computer refers to an information processing device equipped with a control unit, a memory device, etc., and the content monitoring server 1, the customer terminal 4, and the web server 7 are each information processing devices equipped with a control unit, a memory device, etc., and are included in the concept of a computer.
[0030] <Processing Description> Next, the processing related to the content monitoring system 100 will be described. FIG. 2 is a flowchart showing the main processing in the content monitoring server 1 according to this embodiment. FIG. 3 is a flowchart showing the content estimation process in the content monitoring server 1 according to this embodiment. FIG. 4 is a diagram showing an example of HTML data 30 of a web page according to this embodiment. FIG. 5 is a diagram showing an example of image data 50 of a web page according to this embodiment. FIG. 6 is a flowchart showing the page update confirmation process in the content monitoring server 1 according to this embodiment. FIG. 7 is a flowchart showing the content monitoring process in the content monitoring server 1 according to this embodiment.
[0031] The main processing shown in FIG. 2 is started when the processing program is started in the content monitoring server 1 and request information is received from the customer terminal 4 for the first time. In step S (hereinafter, "step S" will be simply referred to as "S") 11, the control unit 10 (request information receiving unit 11) of the content monitoring server 1 determines whether or not request information has been received from the customer terminal 4. If request information has been received (S11: YES), the control unit 10 moves the process to S12. On the other hand, if request information has not been received (S11: NO), the control unit 10 moves the process to S13. Immediately after the start of this main process, this is the first time that request information has been received, so the process becomes YES.
[0032] In S12, the control unit 10 performs a content estimation process. The content estimation process will now be described with reference to FIG. In S21 of FIG. 3, the control unit 10 (content estimation unit 13) estimates the content posted on the web page using the request information. Here, the content estimation method will be explained using a specific example.
[0033] (1) The content is estimated using the HTML data 30 of the web page corresponding to the URL of the request information. FIG. 4 shows an example of some lines extracted from HTML data 30 of a web page requested by customer terminal 4. Here, a case will be described in which the URL of the web page indicated by the HTML data 30 in Fig. 4 and the title "The Story of XXX" are given as request information. In this case, the control unit 10 first searches the HTML data 30 for the title "The Story of XXX". Then, when the character string "XXXXX Story" shown in title 32 is found, the control unit 10 next identifies tag 31 as tag information related to title 32 from the tags before and after title 32 "XXXXX Story." The control unit 10 also identifies tags 33a, 33b, 33c, ... that have a structure similar to tag 31 from HTML data 30, and acquires titles 34 of other content (titles 34a, 34b, 34c, ...) based on each identified tag 33 (tags 33a, 33b, 33c, ...). Then, the control unit 10 causes the storage unit 20 to store information relating to the method by which the content was estimated.
[0034] By carrying out such processing by the control unit 10, if the request information contains at least one title, it is possible to estimate all content, including other content, published on the web page corresponding to the URL using the title and URL. The control unit 10 can then acquire information related to the content, such as the title of the estimated content. Therefore, it is not necessary to include all titles in the request information, which reduces the burden of making a request. In the above-described method, all content included in the HTML data 30 is estimated by identifying the tag 31 as tag information related to the title 32, but the present invention is not limited to this. For example, if a tag is specified in advance, the control unit 10 may extract the title from the tag.
[0035] (2) The content is estimated from the image data 50 of the web page corresponding to the URL of the request information. FIG. 5 shows an example of image data 50 of a web page requested by the customer terminal 4. As shown in FIG. The control unit 10 first identifies each content image 51 (51a, 51b, 51c, ...) from the image data 50 of the web page shown in FIG. 5. The control unit 10 can identify the content image 51, for example, by the size or shape of the content image 51. Next, the control unit 10 performs character recognition on each content image 51 to extract character strings. For example, in the case of content image 51a, the control unit 10 can extract the character strings "aiu" and "ka kaki ku ke ko." Then, the control unit 10 uses the extracted character strings and the title name included in the request information to estimate all content of the web page, including other content. Then, the control unit 10 acquires information related to the content from the estimated content and the acquired character strings. Then, the control unit 10 causes the storage unit 20 to store information relating to the method by which the content was estimated.
[0036] (3) Estimate the content using a learning model. The learning model is learned in advance and stored in the storage unit 20. Furthermore, the learning model can be learned by various methods. The control unit 10 applies a convolutional neural network (CNN) to, for example, image data (not shown) of a web page indicated by HTML data 30 illustrated in FIG. 4 or image data 50 of a web page illustrated in FIG. 5, thereby learning a learning model for estimating content. Then, the control unit 10 uses the learned learning model to estimate content from the image data of a web page corresponding to a URL included in the request information. The control unit 10 also acquires information related to the content of the estimated content.
[0037] As another method, the control unit 10 applies a recurrent neural network (RNN) to the HTML data 30 illustrated in Fig. 4 to learn a learning model for estimating content. Then, the control unit 10 uses the learned learning model to estimate content from the HTML data of a web page corresponding to a URL included in the request information. The control unit 10 also acquires information related to the estimated content.
[0038] In S22 of FIG. 3, the control unit 10 (content estimation unit 13) stores information relating to the acquired content in the content information storage unit 22. In S23, control unit 10 stores a web page corresponding to the URL included in the request information in page storage unit 23. The web page stored in page storage unit 23 may be image data or HTML data. Thereafter, control unit 10 shifts the process to S17 in FIG. 2.
[0039] On the other hand, in S13 of FIG. 2, the control unit 10 performs a page update confirmation process. Here, the page update confirmation process will be described with reference to FIG. In S31 of FIG. 6, the control unit 10 acquires a web page corresponding to a URL stored in the content information storage unit 22. In S32, the control unit 10 compares the content of the web page stored in the page storage unit 23 with the content of the web page acquired in the process of S31. Here, the control unit 10 compares the format of the data (image data or HTML data) stored in the web pages stored in the page storage unit 23. Thereafter, the control unit 10 moves the process to S14 in FIG.
[0040] In S14 of Fig. 2, the control unit 10 (update detection unit 17) determines whether an update has been detected as a result of the comparison made in the process of S32 of Fig. 6. If an update has been detected (S14: YES), the control unit 10 proceeds to S15. On the other hand, if an update has not been detected (S14: NO), the control unit 10 proceeds to S17. In S15, the control unit 10 (content estimation unit 13) estimates the content posted on the updated web page. Then, the control unit 10 (content estimation unit 13) acquires information related to the estimated content and stores the acquired information related to the content in the content information storage unit 22. In S16, the control unit 10 (page update unit 18) updates the web page stored in the page storage unit 23 to the detected updated web page. In S17, the control unit 10 performs a content monitoring process.
[0041] Here, the content monitoring process will be described with reference to FIG. In S41 of FIG. 7, the control unit 10 (crawling unit 14) crawls the websites provided by the web server 7 in order. In S42, the control unit 10 (determination unit 15) determines whether or not the crawling process has found content stored in the content information storage unit 22 on other websites. If the content is found (S42: YES), the control unit 10 proceeds to S43. On the other hand, if the content is not found (S42: NO), the control unit 10 proceeds to S11 in FIG. 2.
[0042] In S43, the control unit 10 (determination unit 15) refers to the content information storage unit 22 and determines whether the collected content is legitimate content. For example, if the file name, location information, content title, etc. of the collected content match those stored in the content information storage unit 22, the control unit 10 determines that the collected content is legitimate content. If the content is legitimate (S43: YES), the control unit 10 shifts the process to S11 in FIG. 2. On the other hand, if the content is not legitimate (S43: NO), the control unit 10 shifts the process to S44.
[0043] In S44, the control unit 10 (notification unit 16) notifies the customer terminal 4 of the URL of the web page on which the content determined to be unauthorized is posted. After that, the control unit 10 shifts the process to S11 in FIG. The control unit 10 then periodically checks whether the web page corresponding to the URL included in the request information has been updated and whether any unauthorized content is posted on the website by repeating the main processing shown in Fig. 2. Furthermore, when the control unit 10 receives another request information from the customer terminal 4, it performs a process of estimating the content posted on the web page corresponding to the URL included in the request information based on the request information.
[0044] As described above, the content monitoring system 100 of this embodiment has the following advantages. (1) The content monitoring server 1 receives request information from the customer terminal 4, the request information including the URL of a web page on which content is posted and the title of the content that identifies the content posted on the web page. Then, the content monitoring server 1 estimates the content posted on the web page based on the request information. Therefore, by receiving request information from a customer that includes a legitimate web page and the title of at least one piece of content posted on that web page, it is possible to estimate the titles, etc., of other pieces of content posted on the web page. As a result, even if a plurality of pieces of content are posted on a web page, it is possible to automatically estimate the pieces of content posted on the web page without specifying all of the pieces of content.
[0045] (2) The content monitoring server 1 refers to the HTML data of the web page, estimates the content from tag information related to the title of the content, and acquires information related to the estimated content. Therefore, by using the HTML data of a web page, it is possible to estimate the titles and the like of all the content, including other content, posted on that web page.
[0046] (3) The content monitoring server 1 performs character recognition processing on the content image of the web page, estimates the content from the relationship between the title of the content and the content image, and acquires information related to the estimated content. Therefore, it is possible to obtain the titles of all the contents, including other contents, posted on a web page from the content image included in the web page.
[0047] (4) The content monitoring server 1 estimates the content posted on the webpage using a learning model. The learning model may be, for example, one that is learned by applying CNN based on the feature quantities of the image data of the webpage, or one that is learned by applying RNN using the strings and tag information connections indicated by the HTML data of the webpage. Therefore, it is possible to estimate content from a web page using a learning model that has previously learned the relationship between web pages and content.
[0048] (5) The content monitoring server 1 stores the web page corresponding to the URL included in the request information in the page storage unit 23, and periodically checks the web page corresponding to the URL to detect updates to the web page stored in the page storage unit 23. When an update is detected, the content monitoring server 1 estimates the content published on the updated web page. Therefore, even if a webpage corresponding to the URL of the webpage included in the request information is updated and, for example, new content is posted on the webpage, the new content can be estimated. Furthermore, for the estimated new content, it can be determined whether new content has been posted on another website. If new content has been posted on another website, it can be determined whether the posted content is legitimate. As a result, once request information is received, even if the content of the webpage corresponding to the URL included in the request information is changed, it is possible to determine whether the changed content is legitimate and notify the URL of the webpage on which the illegitimate content is posted.
[0049] (6) The content monitoring server 1 collects web pages that are posted on other websites of the web server 7 and that contain the estimated content, determines whether the content posted on the collected web pages is legitimate content, and if the content is not legitimate content, notifies the URL of the web page. This allows the system to automatically determine whether the estimated content is posted on other websites, and if so, whether the content is legitimate. Furthermore, if the content is not legitimate, the system notifies the URL of the webpage, allowing for subsequent actions such as warning the administrator of the website posting the illegal content or requesting the removal of the content from the website. Furthermore, crawling websites makes it possible to quickly detect the uploading of illegitimate content.
[0050] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments. Furthermore, the effects described in the embodiments are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiments. Note that the above-described embodiments and the modified embodiments described below can be used in appropriate combinations, but detailed description thereof will be omitted.
[0051] (Variations) (1) In this embodiment, a plurality of methods are exemplified as methods for estimating the content posted on the web page corresponding to the URL included in the request information, but the present invention is not limited to using only one of these methods. For example, the content may be estimated by using a combination of a plurality of methods.
[0052] (2) In the present embodiment, an example has been described in which content posted on other web pages is determined to be legitimate content, but this is not limiting. For example, a specific web page that is not included in the request information but that is approved by the customer may be considered as another web page. Among such web pages, for example, the URL of the specific web page approved by the customer may be stored in a storage unit of the content monitoring server. In this way, the web page stored in the storage unit can be excluded from monitoring, thereby somewhat narrowing the crawling range of the web pages to be monitored.
[0053] (3) In the present embodiment, the content monitoring server is provided with a content information storage unit, a page storage unit, etc., but this is not limiting. A device managed by a customer, such as a customer server (not shown), may be provided with a page storage unit, for example, and the content monitoring server may communicate with the customer's device to reference the page storage unit. Alternatively, a database server or the like may be provided, allowing communication with the content monitoring server, and various types of information may be stored in the database server.
[0054] (4) In the present embodiment, digital content such as e-books and movie content has been exemplified as content, but the content is not limited to this. For example, the content may be a document file consisting of only character data. [Explanation of symbols]
[0055] 1 Content Monitoring Server 4. Customer terminal 7 Web Server 10 Control Unit 11 Request Information Reception Department Page 12 Registration Section 13 Content Estimation Unit 14 Crawling section 15 Judgment section 16 Notification Department 17 Update detection unit 18 Page Update Section 20 Memory section 21 Program memory section 22 Content information storage unit 23 Page Memory 30 HTML data 50 Image data 100 Content Monitoring System
Claims
1. a base value receiving means for receiving an input value that is a content name for identifying the content and a base value including address information of a web page on which the content is posted; a content estimation means for searching for the input value included in the markup description of the web page corresponding to the address information of the basic value received by the basic value receiving means, identifying tag information related to the input value from the description position of the obtained input value, identifying other tag information having a similar structure to the identified tag information, identifying the names of other content based on the identified other tag information, estimating content posted on the web page including other content, and acquiring information related to the content of the estimated content; a collection means for crawling a website where the content can be viewed and / or downloaded, and collecting other web pages on which the content estimated by the content estimation means is posted; a content determination means for determining whether the content of the other web pages collected by the collection means is legitimate content based on whether information related to the content of the other web pages matches the information related to the content acquired by the content estimation means, and for identifying the content of the other web pages that do not match and are determined to be illegitimate content as illegal content to be monitored; A content monitoring device comprising:
2. a base value receiving means for receiving an input value that is a content name for identifying the content and a base value including address information of a web page on which the content is posted; a content estimation means for identifying a content image included in the web page corresponding to the address information of the basic value received by the basic value receiving means, performing character recognition on the identified content image to extract a character string, estimating a content posted on the web page using the input value of the basic value and the character string, and acquiring information related to the content; a collection means for crawling a website where the content can be viewed and / or downloaded, and collecting other web pages on which the content estimated by the content estimation means is posted; a content determination means for determining whether the content of the other web pages collected by the collection means is legitimate content based on whether information related to the content of the other web pages matches the information related to the content acquired by the content estimation means, and for identifying the content of the other web pages that do not match and are determined to be illegitimate content as illegal content to be monitored; A content monitoring device comprising:
3. a base value receiving means for receiving an input value that is a content name for identifying the content and a base value including address information of a web page on which the content is posted; a content estimation means for inputting the web page and the input value corresponding to the address information of the base value accepted by the base value acceptance means into a learning device that has learned to estimate the content of the web page using the web page on which the content is posted and an input value that is the name of the content that identifies the content, and acquiring information related to the content of the content of the web page estimated by the learning device; a collection means for crawling a website where the content can be viewed and / or downloaded, and collecting other web pages on which the content estimated by the content estimation means is posted; a content determination means for determining whether the content of the other web pages collected by the collection means is legitimate content based on whether information related to the content of the other web pages matches the information related to the content acquired by the content estimation means, and for identifying the content of the other web pages that do not match and are determined to be illegitimate content as illegal content to be monitored; A content monitoring device comprising:
4. In the content monitoring device according to claim 3, A content monitoring device, wherein the learning device is a model that has been trained based on features using images of web pages on which the content is posted.
5. In the content monitoring device according to claim 3, The learning device is a model that is trained using connections between character strings and tag information indicated by markup descriptions of web pages on which content is posted.
6. In the content monitoring device according to any one of claims 1 to 5, The content monitoring device further comprises a notification unit that notifies the address information of the other website collected by the collection unit when the content is determined to be unauthorized by the content determination unit.
7. 7. The content monitoring device according to claim 1, a page registration means for storing the web page indicated by the address information of the base value received by the base value receiving means in a page storage unit; update detection means for periodically checking the web page corresponding to the address information of the base value received by the base value receiving means and detecting an update of the web page stored in the page storage unit; a page update means for updating the web page stored in the page storage unit to the updated web page detected by the update detection means; Equipped with The content estimation means estimates content published on the updated web page from the updated web page detected by the update detection means, and acquires information related to the content.
8. A program for causing a computer to function as the content monitoring device according to any one of claims 1 to 7.
Citation Information
Patent Citations
Far infrared heater
JP1979013028A
System for searching illegitimate use of contents
JP2004112318A
Countermeasure system for content illegal distribution
JP2011034251A
Website comparison processing program, website comparison method and device for comparing website
JP2018156198A
Illegal content search device, illegal content search method, and program
JP2018180913A