Pre-release evaluation of digital content
Patent Information
- Application Number
- JP2026513666
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-30
- Filing Date
- 2024-08-29
- Publication Date
- 2026-09-14
Smart Images

Figure 2026531077000001_ABST
Abstract
Description
[[TECHNICAL FIELD]]
[0001]
[0001] The present disclosure relates to the training and use of machine learning evaluation and decision-making models that analyze digital content from various sources and identify inappropriate content before publication. [[BACKGROUND ART]]
[0002]
[0002] Many online platforms use methods to identify, analyze, and remove disinformation generated by third-party (e.g., individual, corporate, AI) users and / or valueless (or incorrect) digital content (inappropriate digital content). These approaches generally rely on post-publication review to assess authenticity. Post-publication review does not properly account for the fact that after inappropriate digital content is published, even if only for a short period of time, it will be viewed and trusted by at least some people. Furthermore, due to the nature of the process, post-publication review often misses inappropriate digital content, and as a result, a high proportion of inappropriate information remains accessible to view. While other methods rely on semantic analysis of digital content prior to publication, such approaches can only detect inappropriate content via very simple operational schemes. These systems cannot detect content generated by more complex inappropriate content generation, such as conversational AI. [[SUMMARY OF THE INVENTION]]
[0003]
[0003] Various embodiments of the present disclosure relate to machine learning methods that may be implemented by a data integrity system operator's computing system. The method includes the step of generating data markers for fraudulent digital content by applying the digital content to a model of known peer data points, where the relationship between the known peer data points and the digital content assesses the likelihood of the digital content being fraudulent. The method may include the step of sending the digital content to a machine learning analysis model to generate an authenticity assessment of the particular digital content. The machine learning assessment model may be trained by applying machine learning to the analysis model and the digital content in order to integrate machine learning to understand a large set of digital content.
[0004]
[0004] Various embodiments of the present disclosure relate to a machine learning computing system comprising a processor and a memory containing instructions that can be executed by the processor. Instructions may include generating data markers for fraudulent digital content by processing digital content with a model of known peer data points and identifying relationships between them and the digital content. The machine learning platform may include sending the digital content to a machine learning evaluation model to generate an evaluation of the fraudulent digital content. The machine learning evaluation model may be trained by applying machine learning to the data points of the digital content and known peer digital content.
[0005]
[0005] Various embodiments of the present disclosure relate to machine learning methods that can be implemented by a data integrity system operator's computing system. These methods may include generating digital content indicators by identifying singularities to develop known peer data points and applying them to identify and evaluate fraudulent digital content. These methods may include sending digital content to a machine learning evaluation model to generate an evaluation of the fraudulent digital content. The machine learning evaluation model may be trained by applying machine learning to network and non-network features.
[0006]
[0006] In various exemplary embodiments, the network model may include consumer reviews. The machine learning platform may define at least one singularity in the consumer reviews based on at least one of the distance between reviews and the time difference between responses, set a fingerprint, and identify patterns between the digital content of the consumer reviews and known peer digital content data points. The machine learning platform converts each consumer review into a metric value to identify patterns and similarities within the set of consumer reviews. These patterns are converted into similarity ratings, and by understanding the common source of content, fraud can be identified.
[0007]
[0007] In various exemplary embodiments, the network model may include commentaries. The machine learning platform converts each commentary into a metric value to identify patterns and similarities within the set of commentaries. These patterns are converted into similarity assessments to understand common sources of information, which can be used to identify fraud.
[0008]
[0008] In various exemplary embodiments, the model may include an online retail site that displays consumer reviews. The machine learning platform can convert each consumer review into a metric value to identify patterns and similarities within the set of consumer reviews. These patterns can be converted into similarity assessments to understand common sources, which can be used to identify fraud.
[0009]
[0009] In various exemplary embodiments, the network model may include a publicly available social comment platform such as a social media platform or a news organization platform. The machine learning platform may convert each social comment into a metric value to identify patterns and similarities within a set of social commentaries. These patterns can be converted into similarity assessments to understand common sources of information, which can then be used to identify fraud.
[0010]
[0010] In various exemplary embodiments, the model may include a distance operator. Known data points may include distance points for determining the distance and / or time between a review and known peer data points. The machine learning platform may be configured to generate false digital content data markers by creating data points from combinations of distance operators and known data points that can be compared.
[0011]
[0011] In various exemplary embodiments, the model may include a network that corresponds to the relationships between known data points. The machine learning platform may convert each relationship into a metric value to identify patterns and similarities within the set of relationships. These patterns can be converted into similarity assessments to understand common sources of information, which can be used to identify fraud.
[0012]
[0012] Various embodiments of the present disclosure relate to machine learning methods implemented by a data integrity system operator's computing system. The method can train one or more evaluation models for evaluating digital content and determine its authenticity before publication. The method may include generating a model of known peer data points and correlating the relationship between digital content and known peer data points. The network model may include at least one digital content unit (e.g., consumer reviews, comment submissions, etc.).
[0013]
[0013] Each singularity within a digital content unit is defined based on at least one of the following: the distance between reviews, the time difference between answers to follow-up credibility questions, the setting of a fingerprint, and the identification of a pattern between the generated digital content and data points of known peer digital content. This method may include generating digital content data points by applying known peer data points to a model. This method may include applying machine learning to known fraudulent peer digital content data points to train a machine learning evaluation model configured to generate evaluations of fraudulent digital content.
[0014]
[0014] Various embodiments of the present disclosure relate to machine learning methods implemented by a computing system for pre-publication analysis and evaluation of digital content, wherein the pre-publication digital content is received and verified by the computing system. A machine learning evaluation model is then operated using known peer data points as training data, and this machine learning evaluation model is trained to generate pre-publication identification of fraudulent digital content. The machine learning evaluation model uses known data points as input to generate evaluations for determining whether the digital content is fraudulent, and the machine learning evaluation model includes density-based clustering techniques, which are functions of density operators. The machine learning evaluation model defines each singularity in the network according to at least one of the quantity, frequency, or occurrence rate of data points between the received digital content and known peer data points. The model includes distance operators, and generating peer data records involves the computing system generating features of distance operator and distance vector combinations compatible with the corresponding distance operators.
[0015]
[0015] Various embodiments of the present disclosure include a machine learning computing system for detecting malicious digital content, the computing system including a processor and memory containing instructions that can be executed by the processor, the instructions including a machine learning platform configured to train a machine learning evaluation model using known peer data points, whole or in part, as training data, the machine learning evaluation model being trained to generate an evaluation of whether the digital content is malicious. An execution by the machine learning evaluation model takes known peer digital content as input and generates an evaluation of whether the digital content is malicious. The model screens the pre-publication digital content at least once for malicious data points, and the machine learning platform defines each singularity in the digital data based on at least one of the frequency or occurrence rate of the known peer data.
[0016]
[0016] The model includes a network that corresponds to the relationships between known peer data records, which contain data points. The machine learning platform defines each singularity in the network as representing a relationship between a corresponding tuple of known peer digital content data points and the pre-publication digital content, and each singularity is weighted according to the characteristics of the corresponding relationship.
[0017]
[0017] These and other features will be described and understood in detail by the following detailed description and accompanying drawings. [Brief explanation of the drawing]
[0018] [Figure 1] Figure 1 is a flowchart showing an overall overview of the system. [Figure 2] Figure 2 is a block diagram of a data integrity system operator / computation system that implements a machine learning platform and communicates with a third-party platform. [Figure 3] Figure 3 is a process flow diagram of a machine learning approach for detecting fraudulent digital content. [Figure 4] Figure 4 is a decision tree diagram illustrating the steps involved in digital content analysis. [Modes for carrying out the invention]
[0019]
[0022] In various embodiments presented herein, machine learning approaches, including supervised and / or unsupervised learning, can be used to train and implement evaluation machine learning models for detecting false or misleading information (i.e., disinformation) before its publication. Throughout, the term false digital content includes, but is not limited to, disinformation, or information that is not true, misleading, or false. False digital content can occur in contexts such as reviews found on shopping, travel, accommodation, restaurants, and restaurant-related sites, or comments on publicly accessible news, social media, and political sites, but is not limited to these. Digital content includes, but is not limited to, user- and / or machine-generated reviews and comments created for publication on websites and platforms on the Internet. False digital content may originate from, but is not limited to, individuals, companies, paid content creators, and / or machine learning / artificial intelligence sources.
[0020]
[0023] This method can be used to identify the relationship between the representation of digital content and known peer groups of digital content data points, and to convert identified digital content data points into benchmarking peer data points. By applying machine learning models, it is possible to analyze the relationships between different data points and understand those relationships, thereby enabling the effective and efficient detection of inappropriate digital content before publication.
[0021]
[0024] A machine learning platform trains and uses a model to leverage the fact that certain digital content may contain data points indicating that it is malicious. Each singularity of digital content can be defined as containing associations between known malicious data points, which may be weighted according to one or more properties of known peer group data points. A singularity indicates an association between a corresponding tuple of known peer digital content data points and the pre-published digital content, where the computational system defines each singularity as indicating an association between corresponding tuples, where a pair is a set of two, and n tuples are a set of n elements of the digital content, and each singularity is weighted according to the properties of the corresponding digital content. This method may include generating malicious data points by applying known peer data points to a model. This method may include applying machine learning to known peer group data points to train a machine learning evaluation model configured to generate evaluations of malicious digital content, and may be described by a remote operator.
[0022]
[0025] In various exemplary embodiments, the method may include comparing digital content data points corresponding to known peer group data points in order to generate an assessment of fraudulent digital content. This process takes place after the creation of the digital content but before publication.
[0023]
[0026] The model may include a curved spatial distance operator. Unlike an orthogonal matrix, which generates a result based only on 90-degree angle vectors and creates a square matrix used to represent a finite graph, the distance operator developed and applied herein enables a wide range of operations by allowing curved inclusion regions, thereby providing an opportunity to recognize additional data points. The use of these distance operators further enhances the ability to interpret digital content data points using a plurality of peer data points and known incorrect data points. The known incorrect data points may include identifying similarity of specific digital content by encoding from a combination of the distance operator and corresponding known peer group data points.
[0024]
[0027] The machine learning platform may employ various machine learning techniques in classification and pattern recognition of digital data points to identify "organic" (as opposed to "fabricated") patterns in data and identify anomalous outlier data points.
[0025]
[0028] Unauthorized digital content often cannot be detected even by analyzing individual digital content alone, and can be detected through the relationships between data points and the associations with data points in peer groups. Whether the data points of a peer group are unauthorized is not important; the patterns / distances between data points provide the information necessary for performing evaluation. Certain activities may conceal unauthorized digital content unless placed in a context that has some identifiable association with data points of a known peer group. Such activities may appear normal when viewed alone. The disclosed approach enables the detection and prevention of unauthorized digital content that could not be detected before publication by conventional methods. Furthermore, the disclosed approach reveals relevant connections within digital content that cannot be detected by human users. As disclosed herein, training an evaluation model using a combination of network-based functions and other functions allows digital content to be considered in context rather than alone, thereby obtaining an accurate evaluation.
[0026]
[0029] Exemplary embodiments of the machine learning model described herein improve computer-related technology by performing functions that cannot be executed by conventional computing systems. Furthermore, the embodiments described herein cannot be implemented by humans. The machine learning model can identify unauthorized digital content that would otherwise go undetected by actively comparing digital content corresponding to known peer group data points and known unauthorized data points. Conventional systems include static databases and definitions, which cannot be configured to acquire, store, and update a large amount of information necessary to identify peer group data points and known unauthorized data points.
[0027]
[0030] Referring to FIG. 1, there is shown a flow diagram illustrating the basic operation of digital content via the system 100 of the present invention, according to a potential embodiment.
[0028]
[0031] System 100 includes a platform computing system 102 (here referred to as a Data Integrity Operator (DIO)). The DIO may be defined as a centralized server that stores peer group data points in a database 116. The platform computing system 102 may belong to an online provider of goods and services, or to a platform that accepts user-generated digital content. For example, many online platforms, such as online shopping platforms, travel-related platforms, and restaurant and restaurant-related platforms, accept user-generated digital content in the form of reviews of goods and services. Other platforms, such as news organizations, social media platforms, and political platforms, accept user-generated comments. While the DIO may include online platforms such as retailers and news organizations, various embodiments envision a centralized server, such as a server farm, that houses peer group data points. The DIO 102 further includes memory that hosts the database 116 for storing data points, and a DIO management module 118 for operating the system. The database 116 and the management module 118 communicate with the DIO 102 at communication points 102a and 102b, respectively.
[0029]
[0032] The components of system 100 can be integrated or operationally coupled to one another, either directly or via a network that enables data exchange, etc. System 100 may include one or more processors, memory, a network interface, and a user interface. The memory can also store data points in the database 116. Interface 106 allows system 100 to communicate with DIO 102 by transmitting and receiving via one or more communication protocols, when system 100 is operationally coupled.
[0030]
[0033] When the platform computing system 102 communicates with the interface 106 (104), the computing system initiates the digital content submission and analysis process. The review collection module 110 organizes how users can submit reviews, and an authenticity check 120 is initiated, collecting information such as submission time, IP address, email address, topic of the digital content, and keywords, but not limited to these. Simultaneously or later, the review collection module 110 sends the digital content to the review collection widget 114 (112). The review collection widget 114 includes a link to the platform 102 website, informing users that they can place / submit digital content there. The authenticity check 120 checks the order of the digital content itself without comparing it with other content or data points. The authenticity check can be performed via email with questions regarding the authenticity of the submitted digital content. The computing system can check the relationship between what is written in the review and the answers to the questions regarding authenticity. Authenticity is determined by analyzing basic characteristics of the digital content, such as consistency in subject and email address. Digital content submissions may be missed or may proceed through this analysis process. In either case, verification check 122 is initiated. For reviews that fail the authenticity check, data points are collected and stored in database 116, making them accessible to machine learning systems for training and use to improve understanding of what constitutes fraudulent content. Subsequent content items can then be compared to the failed items and their characteristics. Verification checks are well-known in the industry. Digital content verification is performed by (1) identifying the IP address of the digital content and (2) sending a confirmation to the sender inquiring about the digital content. This verification is generally done via email, but can be done through any system, such as text, direct mail, or phone.
[0031]
[0034] This invention assigns metrics to data points that previously lacked them. By recording digital content features such as word structure, sentence structure, proximity checks, and feature similarity, along with their time lags (temporal patterns), and combining and comparing these features, it identifies patterns within the digital content that could indicate fraudulent content. As a non-limiting example, if multiple digital content pieces are received and all have the same or similar time lags in their sentence structure, doubts arise about the authenticity of the digital content. Even with different time lags, if a pattern indicating potential fraudulent digital content appears, the machine learning computing system can detect it and store it for future digital content analysis. This is done for all features, individually or in combination. Therefore, this invention is particularly effective in identifying fraudulent digital content generated by artificial intelligence platforms.
[0032]
[0035] Once the initial credibility check is complete, the model performs a primary analysis124, where the data points of each digital content are coded and then compared to other known peer group data points for the credibility of the digital content through measuring the distance between other comparable digital content and consistency within validation, using different known peer group data points. The probability of a digital content being valid is determined by measuring the distance of the data point to the data points of the known peer group. Further judgments are made after multiple iterations (at least 2, and sometimes more than 70).
[0033]
[0036] The model then initiates a further analysis in secondary analysis 126, where patterns between groups of digital content can be identified through measurement. In secondary analysis, patterns of all forms, including time and text, are considered. For example, a machine creating content will generate identifiable patterns. Such analysis is not easily identifiable from a single piece of content and is made possible by examining patterns in peer groups. Again, secondary analysis is performed through multiple iterations of 70 or more.
[0034]
[0037] The analysis results are weighted against known peer group data points, and a digital content verification evaluation 128 is performed. The machine learning model of the present invention is based on the following hypothesis: (i) Unauthorized digital content is fabricated, (ii) The source of the illegal digital content is often the same, (iii) Common sources leave traces in the characteristics of digital content, (iv) The assessment of the likelihood of fraud can be done by detecting traces of a common source in the review in question.
[0035]
[0038] Based on the above, the machine learning model of the present invention includes the following protocol: 1. Digital content collection process: A process for collecting digital content such that (i) the likelihood of authenticity is ensured during the collection stage, and (ii) an array of features for each digital content is generated. 2. Initial assessment of authenticity: Identify the following red flags or white flags regarding digital content: (i) The history of the vendor, product or service under review, (ii) The history of the digital content creator. 3. Measuring the characteristics of the target digital content: Separate and classify individual characteristics and establish a metric operator for each characteristic. 4. Classification of the intrinsic authenticity of each new single digital content: Identify the intrinsic characteristics of each digital content, obtain an initial estimate of its authenticity, and determine the specific analytical module to use. 5. Measuring the distance between the digital content in question and other comparable digital content: Calculate and normalize the distance between single digital content within a specific population and identify patterns between data points of the digital content. 6. Evaluating the authenticity of newly received single digital content: The authenticity of data points in digital content is evaluated by combining the intrinsic plausibility of the single digital content with the distance evaluation of other digital content. 7. Optimize accuracy by introducing machine learning techniques: Improve the accuracy of digital content analysis by strengthening confidence points in digital content and identifying single distance patterns between digital content as suspicious.
[0036]
[0039] Figure 2 shows an example flow of DIO's system 200 interacting with third-party sites that display user content via the III website. The third party 202 is a site or platform that displays user-created digital content. Common sites include retailers, accommodation providers, travel sites, information aggregation sites, news, social media, and political sites. These sites may provide a specific place for user-created digital content (e.g., a retail shopping site that allows product reviews) or may be designed to accept comments in general (e.g., a social media platform) (204). The third party 202 contacts DIO's website via a landing page 208 (206). The third party 202 interfaces with DIO via III, signs up for DIO, and, among other things, submits criteria for digital content (210a). Once the submission of digital content is verified and activated (210b), the third party 202 sends the digital content to DIO for analysis 214 using a machine learning model 216. The mapping of third-party digital content 216a, as described above, includes credibility analysis 216b, primary analysis 216d, and secondary analysis 216d, and is performed by the model using DIO's data point database, which is implemented by a computing system employing a machine learning model (see above description, Figure 1). A determination is made as to whether the digital content is fraudulent (216d). The determination result is sent to third party 202 via the third-party website as feedback 220, 222 (218), and further sent to the DIO-dedicated dashboard 222b. The third party receives feedback from the dashboard and decides internally whether to publish the digital content or uses a decision proposal generated by the computing system.
[0037]
[0040] Figure 3 shows a process flow diagram of a machine learning approach for detecting fraudulent digital content. The machine learning system 300 acquires the content to be checked 302 and performs an authenticity check of the digital content 304 as described above with reference to Figure 1. The digital content is sent to the digital content synthesis 308 (306), where it is coded into dimensions defined for further analysis. The machine learning system 300 then sends the synthesized digital content to the analysis and mapping 312 (310). The digital content is decomposed into a single review vector, i.e., data points (312a). The single review vector is then mapped to a DIO numerical review model, i.e., known peer group data points (312b), and a pattern analysis 312c is performed between the single digital content vectors, benchmarking them against known peer group content. Once the pattern and benchmark analysis is complete, the machine learning system 300 analyzes the results (316) to evaluate the authenticity of the digital content (316a) and to determine whether the digital content should be published (316b).
[0038]
[0041] Subsequently, the machine learning system 300 provides evaluation feedback to the peer group database to train the model (316c), and at the same time provides feedback to the requesting website (316d).
[0039]
[0042] Figure 4 shows an example of a decision tree for a machine learning model. At 402, the user submits digital content they wish to publish. The model collects predetermined information, which may include an email address, IP address, and text (including keywords, title, product or service identifiers, etc.). At 404, an initial analysis is performed to determine if the submission is from a verified user. If so (408), the digital content is published (410), and the content's features are stored in database 116 (Figure 1) for the machine learning model's training and subsequent analytical integration (412). If the digital content is not from a verified user (413), further analysis 414 is performed as described above. A credibility question is sent to the user (416). If the credibility question is not answered (418), the digital content is blocked.
[0040] If no answer is provided to the credibility question, the content features are not saved. If an answer is provided to the credibility question (422), the computing system performs an analysis to determine whether the provided answer is valid. Otherwise (424), the digital content is blocked, and the content features are saved in database 116 (Figure 1) for the purpose of training a machine learning model and for subsequent analytical integration (426). If the digital content is credible (428), further analysis 430 is performed, and the content features are converted into data points and matched against known peer group data points (i.e., primary analysis).
[0041]
[0043] If the digital content features (data points) are similar to known data points from a peer group's authentic review (444), or similar to data points from known fraudulent digital content (446), the digital content is blocked, and the content features are stored in database 116 (Figure 1) for the purpose of training a machine learning model and for subsequent analytical integration (447). If the digital content data points are different from authentic known peer group data points or known fraudulent data points (432), the digital content data points are further checked for content features within the known peer group (434) (i.e., secondary analysis). If the digital content features are not part of a peer group pattern of authentic or fraudulent digital content (436), it is exposed (438), and the content features are stored in database 116 (Figure 1) for the purpose of training a machine learning model and for subsequent analytical integration. If the characteristics of a digital content data point are similar to a known pattern of genuine data point (440), or similar to a known pattern of data point of fraudulent digital content (442), the digital content is blocked, and the content characteristics are stored in database 116 (Figure 1) for the purpose of training a machine learning model and for subsequent analytical integration (446).
[0042]
[0044] Steps 432 and 434 of the machine learning process involve multiple iterations. Typically, there are 70 to 200 data points for each piece of peer digital content. This count is added each time a new post is made, and all new posts are reviewed against peer reviews from known groups. Therefore, each new digital content post may be analyzed 70 to 200 times in each of these two steps, multiplied by the number of reviews from the peer group. While the range of 70 to 200 is common, it is not limiting.
[0043]
[0045] The embodiments described herein are explained with reference to the drawings. The drawings illustrate specific details of specific embodiments that provide the systems, methods, and programs described herein. However, the use of drawings to describe embodiments should not be construed as imposing any limitations on this disclosure that may exist in the drawings.
[0044]
[0046] The embodiments described above are presented for illustrative and explanatory purposes only. This disclosure is not intended to be exhaustive or to limit the disclosure to the exact form disclosed, and modifications and variations are possible in light of the teachings above, or may be obtained from this disclosure. These embodiments are selected and described to illustrate the principles and practical applications of this disclosure so that those skilled in the art can utilize various embodiments and make various modifications.
[0045]
[0047] An exemplary arithmetic system and device may include one or more processing units, each having one or more processors; one or more memory units, each having one or more memory devices; and one or more system buses connecting various components, including the memory units, to the processing units. Each memory device may include non-transient volatile storage media, non-volatile storage media, non-transient storage media (e.g., one or more volatile memories and / or non-volatile memories), etc. In this regard, machine-executable instructions include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a particular function or set of functions. Each memory device may be operable to maintain or otherwise store information relating to operations performed by one or more associated modules, units, and / or engines, including processor instructions and associated data (e.g., database components, object code components, script components, etc.), according to the exemplary embodiments described herein.
[0046]
[0048] While the drawings in this book illustrate a specific sequence and configuration of method steps, it should be noted that the order of these steps may differ from those shown. For example, two or more steps may be performed simultaneously or partially simultaneously. Furthermore, several method steps performed as individual steps may be combined, steps performed as combined steps may be separated into individual steps, the order of a particular process may be reversed or otherwise altered, and the nature or number of individual processes may be changed or altered. The order or arrangement of any element or apparatus may be modified or substituted according to a different embodiment. Therefore, all such modifications are intended to fall within the scope of this disclosure as defined in the appended claims. Such variations will depend on the selected machine-readable media and hardware system, and the designer's choice. All such variations are understood to fall within the scope of this disclosure. Similarly, software and web implementations of this disclosure can be implemented using standard programming techniques with rule-based logic and other logic to perform various database retrieval, correlation, comparison, and decision steps.
[0047]
[0049] The above description of embodiments is presented for illustrative and explanatory purposes only. This disclosure is not intended to be exhaustive or to limit the disclosure to the exact form disclosed, and modifications and variations are possible in light of the above teachings, or may be obtained from this disclosure. These embodiments are selected and described to illustrate the principles and practical applications of this disclosure so that those skilled in the art can utilize the various embodiments and make various modifications to suit specific intended uses. Other substitutions, modifications, changes and omissions may be made in the design, operating conditions and arrangement of embodiments without departing from the scope of this disclosure as set forth in the appended claims.
Claims
1. A machine learning method implemented by a computing system for pre-publication evaluation of illicit digital content, the method comprising: a step of receiving pre-publication digital content by the computing system; a step of verifying the digital content; a step of training a machine learning evaluation model by the computing system using known peer group data points as training data, the machine learning evaluation model being trained to generate pre-publication identification of illicit digital content; a step of generating a model by the computing system that includes a set of trained data point patterns and relationships thereto known to generate illicit digital content, based on a plurality of data records accessed via an electronic database; and a step of pre-publication evaluation of illicit digital content by the computing system. A machine learning method comprising: a step of searching a plurality of known peer group data points from an electronic database; a step of generating an evaluation analysis of fraudulent digital content by screening the peer group data records for digital content using the computing system, wherein the screening is repeated according to a plurality of iterative steps; a step of running the machine learning evaluation model using the known data points as input to generate an evaluation of fraudulent digital content using the computing system, wherein the machine learning model includes density-based clustering techniques which are a function of density parameters; and a step of generating markers for identifying fraudulent digital content using the computing system.
2. The machine learning method according to claim 1, wherein the digital content is selected from a group including consumer reviews, political comments, and healthcare information.
3. The machine learning method according to claim 2, wherein the calculation system defines each singularity in the network according to at least one of the time pattern, quantity, frequency, or occurrence rate of data points between the received digital content and peer group data.
4. The machine learning method according to claim 3, wherein the step of generating the model includes the steps of detecting known fraudulent data points in the peer group data using the computing system, and screening the data points for similarity against the digital content.
5. The machine learning method according to claim 4, wherein digital content containing potentially malicious data points is marked for deletion.
6. The machine learning method according to claim 5, wherein newly identified fraudulent data points are included in the peer group data record.
7. The machine learning method according to claim 6, wherein the step of generating the model includes detecting common relationships for evaluating sources of digital content using the computing system.
8. The machine learning method according to claim 1, wherein the model includes content similarity between digital contents, and further includes a time lag from receiving the digital content to verifying the digital content.
9. The machine learning method according to claim 8, wherein the calculation system defines each singularity as indicating a relationship between corresponding tuples, a pair is a set of two, n tuples are a set of n elements of digital content, and each singularity is weighted according to the characteristics of the corresponding digital content.
10. The machine learning method according to claim 9, wherein the step of generating the model includes a step of using a computing system to detect relevance using stored peer group data records and data accessed via an internet-based platform that publishes digital content, in order to identify similarities between digital content.
11. The machine learning method according to claim 1, wherein the machine learning platform defines each singularity as a content vector containing malicious content.
12. The machine learning method according to claim 11, wherein the model includes a distance operator, and the step of generating the peer group data record includes generating features of a combination of the distance operator and data points compatible with the corresponding distance operator by the calculation system.
13. A machine learning computing system for detecting fraudulent digital content, the computing system comprising a processor and a memory containing instructions executable by the processor, the instructions comprising: a step of training a machine learning evaluation model using all or part of peer group digital content as training data, the machine learning evaluation model being trained to generate evaluations of fraudulent digital content and to generate markers to identify potentially fraudulent digital content; a step of generating a model based on a plurality of data records, including known fraudulent digital content and the relationships between them; a step of running the machine learning evaluation model using peer group digital content as input to generate evaluations of fraudulent digital data, the machine learning evaluation model comprising density-based clustering techniques which are a function of density parameters; and a step of generating markers in accordance with the evaluations of fraudulent digital content.
14. The machine learning system according to claim 13, wherein the model includes pre-public digital content that has been screened at least once for fraudulent data points, and the machine learning platform defines each singularity in the digital data according to at least one of the frequency or occurrence rate of known fraudulent peer group data.
15. The machine learning system according to claim 13, wherein the digital content is selected from a group including consumer reviews, political comments, and healthcare information.
16. The machine learning system according to claim 13, wherein the model includes relationships between records of peer group data points and digital content data points, the machine learning platform defines each singularity as representing a relationship between corresponding pairs of peer group digital content and pre-publication digital content, and each singularity is weighted according to the characteristics of the corresponding relationship.
17. A machine learning computing system for detecting fraudulent digital content, the computing system comprising a processor and a memory containing instructions executable by the processor, the instructions comprising: a step of training a machine learning evaluation model using all or part of peer group digital content as training data, the machine learning evaluation model being trained to generate evaluations of fraudulent digital content and to generate markers to identify potentially fraudulent digital content; a step of generating a model based on a plurality of data records, which includes known fraudulent digital content and the relationships between them; a step of running the machine learning evaluation model using peer group digital content as input to generate evaluations of fraudulent digital data, the machine learning evaluation model comprising a density-based clustering technique which is a function of a density parameter using a distance operator; and a step of generating markers in accordance with the evaluations of fraudulent digital content, the machine learning computing system comprising a machine learning platform configured to do the following:
18. The model further includes authentic digital content, according to claim 17, for the machine learning computing system.
19. The machine learning system according to claim 17, wherein the machine learning evaluation model encodes the data points of each digital content and performs a primary analysis in which the authenticity of the digital content is compared with other known peer group data points over multiple iterations using known peer group data points, and the machine learning evaluation model then performs a secondary analysis in which patterns between groups of digital content are identified through the detection of patterns between other comparable digital content over multiple iterations, and the probability of the digital content being valid is determined by measuring the distance of the data points to known peer group digital content.
20. The machine learning evaluation model according to claim 19, wherein each digital content is coded into a data point and analyzed against peer reviews of all known groups, and each submission of digital content is repeatedly analyzed against the known data point as many times as there are peer group reviews.