Method, device and computer-readable recording medium for learning and processing to determine whether a sentence has violence

KR103014139B1Active Publication Date: 2026-09-04SALTLUX
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020230113342
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-08-29
Publication Date
2026-09-04
Estimated Expiration
2043-08-29

Smart Images

  • Figure 112023095018114-PAT00001_ABST
    Figure 112023095018114-PAT00001_ABST
Patent Text Reader

Abstract

The present invention relates to a learning and processing method for determining the violent nature of a sentence, implemented by a computing device comprising one or more processors and one or more memories for storing instructions executable by said processors, comprising: a vector information generation step in which, when a plurality of natural languages ​​are input by a user, each learning sentence corresponding to said plurality of natural languages ​​is analyzed through a previously stored vector analysis algorithm, and a vector value for each of said learning sentences is calculated through the analysis result to generate a plurality of vector information; and an algorithm learning completion step in which, when the generation of the plurality of vector information is completed, each vector value of said plurality of vector information is identified, and a grouping process is performed for each of said identified vector values, wherein the grouping process is performed according to a plurality of intent categories to complete the learning of an intent classification algorithm. The method is characterized by including a sentence violence determination step, wherein, when the algorithm learning completion step is completed and a target natural language for determining violence is input, the target sentence corresponding to the target natural language is analyzed based on the intent classification algorithm to calculate the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the plurality of intent categories, and then calculates a match value based on whether there is a match between the keywords constituting the target sentence and the keywords constituting the similar sentence, and compares the violence score of the target sentence calculated based on the similarity and the match value with a preset threshold value to determine whether the target sentence is violent. In addition to this, various embodiments identified through this document are possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a learning and processing method for determining whether a sentence is violent. Specifically, the invention relates to a technology that analyzes multiple natural language inputs from a user to identify vector values, performs a grouping process on the identified vector values ​​to classify multiple natural language inputs by multiple intent categories, thereby completing the learning of an intent classification algorithm, and subsequently analyzes the vector values ​​of a target sentence input through the intent classification algorithm to determine whether the target sentence is violent. Background Technology

[0002] Natural Language Processing (NLP) is a technology that enables computers to analyze and understand the syntax and semantics of human language. Word embedding technology, a type of NLP technique, is derived from neural network-based language models and allows for the representation of lexical meaning by arranging similar words closely in a vector space. As this embedding technology advances, new technologies incorporating it are being developed across various fields. In particular, the need for technologies to address issues is increasing in Korea due to the frequent occurrence of problems such as school violence and misconduct within military units. However, these technologies are limited to resolving online issues; consequently, they face difficulties and have clear limitations in practically resolving offline problems, such as school violence and misconduct in the military. Accordingly, the industry is developing various technologies to resolve these aforementioned problems.

[0003] As an example, Korean published patent 10-2021-0101038 (Method for recognizing violent and non-violent situations based on Korean conversation using a BERT language model) discloses a technology for recognizing violent and non-violent situations by analyzing sentences containing Korean-based violent and non-violent situation words and phrases through a BERT language model.

[0004] However, the aforementioned prior art discloses only a technology that analyzes words and phrases within a sentence to identify whether the sentence pertains to either a violent or non-violent situation. It does not disclose a technology that analyzes multiple natural languages ​​input by a user to identify vector values, performs a grouping process on the identified vector values ​​to classify multiple natural languages ​​by multiple intent categories, completes the training of an intent classification algorithm, and then analyzes the vector values ​​of a target sentence input thereafter using the intent classification algorithm to determine whether the target sentence is violent. Consequently, there is a growing need for a technology capable of addressing this issue. The problem to be solved

[0005] Accordingly, the present invention aims to solve and prevent problems occurring offline, such as school violence incidents and irregularities in military units, by applying a technology that identifies whether a sentence contains violence in various fields, through a learning and processing method for determining whether a sentence contains violence, by analyzing multiple natural languages ​​input by a user to identify vector values, performing a grouping process for the identified vector values ​​to classify multiple natural languages ​​by multiple intention categories, and then analyzing the vector values ​​of a target sentence input thereafter through the intention classification algorithm to determine whether the target sentence contains violence. means of solving the problem

[0006] A learning and processing method for determining whether a sentence is violent, implemented by a computing device comprising one or more processors and one or more memories for storing instructions executable by said processors according to an embodiment of the present invention, wherein when a plurality of natural languages ​​are input from a user, a vector information generation step of analyzing each of the learning sentences corresponding to said plurality of natural languages ​​through a previously stored vector analysis algorithm, and generating a plurality of vector information by calculating a vector value for each of said learning sentences through the analysis result; and an algorithm learning completion step of, when the generation of the plurality of vector information is completed, identifying the vector value of each of said plurality of vector information, and then performing a grouping process for each of said identified vector values, wherein the grouping process is performed according to a plurality of intent categories to complete the learning of an intent classification algorithm. The method is characterized by including a sentence violence determination step in which, when the algorithm learning completion step is completed and a target natural language for determining violence is input, the target sentence corresponding to the target natural language is analyzed based on the intent classification algorithm to calculate the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the plurality of intent categories, and then calculates a match value based on whether there is a match between the keyword constituting the target sentence and the keyword constituting the similar sentence, and compares the violence score of the target sentence calculated based on the similarity and the match value with a preset threshold value to determine whether the target sentence is violent.

[0007] The vector information generation step preferably includes: a sentence classification step in which, while a plurality of natural languages ​​are input by a user, the plurality of natural languages ​​are classified according to the plurality of intent categories through input separately received from the user; and a vector calculation step in which, when the sentence classification step is completed, a learning sentence corresponding to the plurality of natural languages ​​is analyzed through the previously stored vector analysis algorithm to calculate an initial vector value based on the meaning of each of the plurality of learning sentences and generate vector information for each of the plurality of learning sentences.

[0008] The algorithm learning completion step described above may include: a grouping completion step in which, when the generation of vector information for each of the plurality of learning sentences is completed, the first initial vector value of the first learning sentence classified into the first intention category among the plurality of intention categories is corrected to a central intention vector value derived based on the vector value of a reference sentence included in the first intention category, thereby completing the grouping of the first learning sentence for the first intention category; and a learning completion step in which, when the grouping process of the plurality of learning sentences for each of the plurality of intention categories is completed as the function of the grouping completion step is repeated multiple times, the learning of the intention classification algorithm is completed.

[0009] The above central intention vector value is configured to be calculated through a vector value corresponding to each reference sentence among the reference sentences included in each of the above multiple intention categories, wherein the usage frequency is learned to be greater than a specified frequency. This configuration is capable of correcting the initial vector value of the above-mentioned training sentence to establish a standard for distinguishing between the above multiple intention categories, thereby preventing cases where reference sentences serving as a standard for determining the violent nature of a target sentence corresponding to the target natural language input in the sentence violence judgment step are not distinguished by the above multiple intention categories, and ensuring that no reference sentences are omitted.

[0010] The above-mentioned plurality of intention categories includes a positive intention category and a negative intention category distinguished based on whether the reference sentence has violent attributes, and is configured such that the central intention vector value is matched to each of the above-mentioned positive intention category and the above-mentioned negative intention category, and includes a reference sentence having a vector value for deriving the central intention vector value, wherein when the learning of the intention classification algorithm is completed according to the performance of the function of the algorithm learning completion stage, the learning sentences corresponding to the above-mentioned initial vector value can be classified as the reference sentence and included in the positive intention category and the negative intention category.

[0011] The sentence violence determination step described above may include: a first vector value calculation step, wherein, when the target natural language is input while the algorithm learning completion step is complete, a target sentence corresponding to the target natural language is analyzed through the intent classification algorithm to calculate a first vector value based on the meaning of the target sentence; a similar intent category identification step, wherein, when the calculation of the first vector value is completed, a similarity calculation process is performed between the calculated first vector value and the vector value of a reference sentence included in the plurality of intent categories, and the intention category containing the similar sentence is identified as a similar intent category along with a similar sentence which is a reference sentence having the highest vector value with respect to the first vector value among the reference sentences included in each of the plurality of intent categories; and a first violence score calculation step, wherein a first violence score is calculated based on the similarity calculated according to the function of the similar intent category identification step and a preset first weight, and is used to determine whether the target sentence is violent.

[0012] The sentence violence determination step may further include: a match value calculation step, which, when the identification of the similar intent category is completed, performs a process to determine whether the keywords constituting the similar sentence included in the similar intent category and the keywords constituting the target sentence are identical, and calculates a match value calculated according to whether each keyword matches; a second violence score calculation step, which, when the calculation of the match value is completed, calculates a second violence score, which is a component used to determine the violence of the target sentence, based on the calculated match value and a preset second weight; and a violence information provision step, which, when the calculation of the first violence score and the second violence score is completed, calculates a sum score by summing the first violence score and the second violence score, checks whether the calculated sum score exceeds a preset threshold value matched to the similar intent category, and if it is confirmed through the check result that the sum score exceeds the preset threshold value, determines that violence exists in the target sentence, and generates sentence violence information through a history based on the function execution of the intention classification algorithm and provides it to the user.

[0013] The above sentence violence information is information generated through a history based on the performance of the function of the above intent classification algorithm, and may include the above target sentence, whether the above target sentence is violent, the intent category to which the above target sentence is included, the reference sentence with the highest similarity to the above target sentence among the reference sentences included in the above intent category, and keywords containing a violent meaning among the keywords included in the above target sentence.

[0014] A learning and processing device for determining the violent nature of a sentence, implemented as a computing device comprising one or more processors and one or more memories for storing instructions executable by said processors according to an embodiment of the present invention, wherein when a plurality of natural languages ​​are input by a user, a vector information generation unit analyzes each of the learning sentences corresponding to said plurality of natural languages ​​through a previously stored vector analysis algorithm, calculates a vector value for each of said learning sentences through the analysis result, and generates a plurality of vector information; and an algorithm learning completion unit, wherein when the generation of the plurality of vector information is completed, identifies the vector value of each of said plurality of vector information, performs a grouping process for each of said identified vector values, and performs the grouping process according to a plurality of intent categories to complete the learning of an intent classification algorithm. The present invention is characterized by including a sentence violence determination unit that, when a target natural language for determining violence is input while the function of the algorithm learning completion unit is completed, analyzes a target sentence corresponding to the target natural language based on the intent classification algorithm, calculates the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the plurality of intent categories, calculates a match value based on whether there is a match between the keyword constituting the target sentence and the keyword constituting the similar sentence, and determines whether there is violence in the target sentence by comparing the violence score of the target sentence calculated based on the similarity and the match value with a preset threshold value.

[0015] A computer-readable recording medium according to an embodiment of the present invention, wherein the computer-readable recording medium stores instructions for a computing device to perform the following steps, the steps comprising: a vector information generation step in which, when a plurality of natural languages ​​are input by a user, each learning sentence corresponding to the plurality of natural languages ​​is analyzed through a previously stored vector analysis algorithm, and a vector value for each of the learning sentences is calculated through the analysis result to generate a plurality of vector information; and an algorithm learning completion step in which, when the generation of the plurality of vector information is completed, each vector value of the plurality of vector information is identified, and a grouping process is performed for each of the identified vector values, wherein the grouping process is performed for each of the plurality of intent categories to complete the learning of the intent classification algorithm. The method is characterized by including a sentence violence determination step in which, when the algorithm learning completion step is completed and a target natural language for determining violence is input, the target sentence corresponding to the target natural language is analyzed based on the intent classification algorithm to calculate the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the plurality of intent categories, and then calculates a match value based on whether there is a match between the keyword constituting the target sentence and the keyword constituting the similar sentence, and compares the violence score of the target sentence calculated based on the similarity and the match value with a preset threshold value to determine whether the target sentence is violent. Effects of the invention

[0016] The learning and processing method for determining the presence of violence according to the present invention determines and learns the presence of violence in a sentence based on vector values ​​based on the meaning of the sentence, and can accurately identify the presence of violence in sentences utilizing various sentence structures and keywords.

[0017] In addition, it can be integrated with technology in other fields and utilized to resolve and prevent issues such as school violence, irregularities in military units, and workplace bullying. Brief explanation of the drawing

[0018] FIG. 1 is a flowchart illustrating a learning and processing method for determining whether a sentence is violent according to an embodiment of the present invention. FIG. 2 is a flowchart illustrating the vector information generation step of a learning and processing method for determining whether a sentence is violent according to an embodiment of the present invention. FIG. 3 is a block diagram illustrating the algorithm learning completion section of a learning and processing system for determining whether a sentence is violent according to an embodiment of the present invention. FIG. 4 is a flowchart illustrating the sentence violence determination step of a learning and processing method for determining whether a sentence is violent according to an embodiment of the present invention. FIG. 5 is a flowchart illustrating the sentence violence determination step of a learning and processing method for determining whether a sentence is violent according to an embodiment of the present invention. FIG. 6 is a drawing for explaining an example of the internal configuration of a computing device according to an embodiment of the present invention. Specific details for implementing the invention

[0019] Hereinafter, various embodiments and / or aspects are disclosed with reference to the drawings. For illustrative purposes, numerous specific details are disclosed in the following description to aid in a general understanding of one or more aspects. However, it will also be recognized by those skilled in the art that these aspects may be practiced without such specific details. The following description and the accompanying drawings describe specific exemplary aspects of one or more aspects in detail. However, these aspects are exemplary, and some of the various methods in the principles of the various aspects may be used, and the description is intended to include all such aspects and their equivalents.

[0020] As used herein, terms such as "examples," "examples," "aspects," "examples," etc., may not be interpreted as implying that any aspect or design described is better or more advantageous than other aspects or designs.

[0021] Additionally, the terms “comprising” and / or “comprising” should be understood to mean that the relevant feature and / or component is present, but not to exclude the presence or addition of one or more other features, components and / or groups thereof.

[0022] Additionally, terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0023] Furthermore, in the embodiments of the present invention, all terms used herein, including technical or scientific terms, unless otherwise defined, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in the embodiments of the present invention.

[0024] FIG. 1 is a flowchart illustrating a learning and processing method for determining whether a sentence is violent according to an embodiment of the present invention.

[0025] Referring to FIG. 1, a learning and processing method for determining whether a sentence is violent, implemented as a computing device comprising one or more processors and one or more memories that store instructions that can be executed on said processors, may include a vector information generation step (step S101), an algorithm learning completion step (step S103), and a sentence violence determination step (step S105).

[0026] In step S101, when multiple natural languages ​​are input from a user, the one or more processors (hereinafter referred to as processors) can analyze each of the learning sentences corresponding to the multiple natural languages ​​through a stored vector analysis algorithm, and generate multiple vector information by calculating a vector value for each of the learning sentences through the analysis results.

[0027] According to one embodiment, the previously stored vector analysis algorithm may be an algorithm that vectorizes learning sentences corresponding to multiple natural languages ​​received from a user so that the processor can understand natural language, which is human language, and calculates a vector value for each of the learning sentences.

[0028] In relation to the above, the stored vector analysis algorithm may be an algorithm including at least one of a one-hot encoding-based model included in a sparse representation category, a word embedding included in a dense representation category, a frequency-based embedding, a bag of words (BOW), a count vector, a Tf-idf (term frequency - inverse document frequency), a prediction-based embedding, Word2Vec, a continuous bag of words (CBOW), and a Skip-gram-based model.

[0029] According to one embodiment, the processor can analyze a learning sentence corresponding to a plurality of natural languages ​​through a previously stored vector analysis algorithm, calculate a vector value of the learning sentence corresponding to each of the plurality of natural languages ​​through the analysis result, and then generate vector information based on the calculated vector value.

[0030] That is, the plurality of vector information may be information containing vector values ​​for a learning sentence corresponding to each of the plurality of natural languages.

[0031] According to one embodiment, when the generation of the plurality of vector information is completed, the processor may perform an algorithm learning completion step (step S103).

[0032] In step S103, when the generation of the plurality of vector information is completed, the processor identifies the vector value of each of the plurality of vector information and performs a grouping process for each of the identified vector values, and can complete the learning of the intent classification algorithm by performing the grouping process for each of the plurality of intent categories.

[0033] According to one embodiment, the grouping process may be a process for clearly classifying the training sentences by multiple intention categories through vector values ​​based on the multiple vector information. More specifically, the grouping process may be a process for analyzing and identifying whether the training sentences are violent, and then clearly distinguishing and classifying criteria regarding the violence of the training sentences through the multiple intention categories, wherein the training sentences having violence are grouped together and the training sentences not having violence are grouped together.

[0034] According to one embodiment, the plurality of intent categories may be configured to distinguish and classify learning sentences based on whether they are violent.

[0035] In relation to the above, the plurality of intention categories may include positive intention categories and negative intention categories. For example, among the plurality of positive intention categories, the first positive intention category (category of anti-bullying attributes) may be a category containing sentences that prevent workplace bullying, and among the plurality of negative intention categories, the first negative intention category (category of bullying occurrence attributes) may be a category containing sentences that encourage bullying behavior involving the ostracization of specific individuals in the workplace.

[0036] More specifically, the plurality of intent categories may include a reference sentence having a vector value for deriving the central intent vector value, configured such that the central intent vector value is matched to each of the positive intent category and the negative intent category. In relation to the above, refer to FIG. 3 for a detailed explanation of the central intent vector value.

[0037] In addition, when the learning of the intent classification algorithm is completed in accordance with the function of the algorithm learning completion step (step S103), the learning sentences corresponding to the initial vector values ​​may be classified as reference sentences and included in the positive intent category and the negative intent category.

[0038] Accordingly, the processor can complete the learning of an intent classification algorithm that algorithmizes the series of processes by performing a grouping process for each vector value of the plurality of vector information and performing the process of classifying multiple training sentences by multiple intent categories multiple times.

[0039] That is, the above intent classification algorithm may be an algorithm for analyzing the vector value of a target sentence input thereafter to determine whether the target sentence is similar to a reference sentence included in any one of the plurality of intent categories.

[0040] According to another embodiment, the intent classification algorithm may include an encoder-based BERT model. In this regard, the BERT model is a pre-trained NLP (natural language processing) technology developed by Google, and it is a general-purpose language model that performs well in all natural language processing fields, rather than being limited to a specific field. Since BERT demonstrates state-of-the-art performance in more than 11 natural language processing tasks, it may be a model that is currently receiving attention as a highly suitable model for language processing.

[0041] If there is a sufficient amount of data, embeddings have a significant impact on the performance of a model for a specific task, and embedded words represented by vectors that accurately express the meaning of the words will naturally yield good performance during the training process. BERT is used in this embedding process, and it can be understood as a language model that can improve the performance of a specific task through pre-trained embeddings before undertaking the task.

[0042] Before the advent of BERT, Word2Vec, GloVe, and Fasttext methods were widely used for data preprocessing embeddings, but nowadays, most high-performance models use BERT.

[0043] In the BERT model, input text data is processed through Token Embedding, Segment Embedding, and Position Embedding. Token Embedding is a Word Piece embedding method that embeds each character individually and forms a single unit from the longest, most frequently occurring sub-word. Infrequently occurring words are broken down into sub-words. This can also resolve the 'OOV' problem, which previously degraded modeling performance by treating all infrequently occurring words as 'OOV'.

[0044] Segment Embedding performs the task of reconstructing tokenized words into a single sentence. In BERT, two sentences are separated using a delimiter ([SEP]) and designated as a single segment for input. BERT limits the length of this segment to 512 sub-words; however, since Korean typically consists of 20 sub-words and most sentences do not exceed 60 sub-words, it is sufficient for training when using BERT by limiting the length of a single segment to 128.

[0045] In position embedding, BERT uses only a part of the Transformer model. A Transformer is a model that uses a model called Self-Attention instead of models like CNN or RNN. BERT uses only the encoder among the encoder and decoder of the Transformer.

[0046] Self-Attention considers the positional information of input tokens rather than the position of the input. Therefore, Transformer models use positional encoding via the sinusoid function, and BERT adopts this approach to use positional encoding. In other words, it performs embedding in the order of the tokens.

[0047] BERT combines the three embeddings mentioned above, applies Layer Normalization and Dropout to them, and uses them as input.

[0048] Once all the data to be trained has been encoded by embedding the data, pre-training is performed. That is, in the present invention, the processor applies a BERT model, and the BERT model does not use an existing analysis model for a specific field, but first presents a general model and then fine-tunes it to derive an algorithm that performs the corresponding task. In other words, pre-training is performed on the BERT model, and the BERT model is continuously trained through fine-tuning using sentences other than the aforementioned sentence.

[0049] Existing methods typically predict the next word by learning sentences from left to right, or by considering the context to the left and right of the word to be predicted. However, BERT uses two methods—MLM (Masked Language Model) and NSP (Next Sentence Prediction)—to effectively learn the characteristics of language.

[0050] MLM is trained by randomly discarding tokens from input sentences (masking) and predicting those tokens, while NSP is a method that predicts the order of two sentences given them; it is trained to predict the association between two sentences to fine-tune NLI and QA, where the relationship between sentences must be considered.

[0051] After pre-training, a fine-tuning process is performed to transfer the language model trained as described above. In other words, the pre-trained word embeddings (sentence embeddings) sufficiently contain the semantic and grammatical information of the corpus, and through additional fine-tuning training for downstream tasks, the embeddings are updated to suit the downstream tasks. Transfer learning is created by stacking additional models on the output of BERT's language model. It is generally known that good performance is achieved even when stacking only a simple DNN, without stacking complex CNNs, LSTMs, or Attention systems, and there is little difference in performance.

[0052] To use BERT for each task, first, for the sentence pair classification problem, two sentences are input as a single input to find the relationship between the two sentences, one sentence is input to classify the type of sentence, the start and end of the desired correct answer position within a sentence or paragraph are found, or the named entity recognition of input sentence tokens is found or the part-of-speech tagging is found.

[0053] At this time, the processor performs positive / negative classification of the target sentence using the corresponding fine-tuning. In the case of a task that classifies the type (positive / negative) of the target sentence using fine-tuning, a classification layer is added to classify positive / negative based on keywords such as numerical values ​​or words in the sentence, and parameters such as Batch Size, Epoch, Max Sequence Length, Dropout Rate, and Learning Rate can be set and used.

[0054] According to another embodiment, the intent classification algorithm may include one of a feature extraction modeling-based algorithm, an embedding vector modeling-based algorithm, and a deep learning modeling-based algorithm.

[0055] In relation to the above, when the processor starts an analysis of a target sentence through the previously stored intent classification algorithm, it can analyze the target sentence through an algorithm based on feature extraction modeling included in the previously stored intent classification algorithm.

[0056] At this time, the processor can identify the morphemes constituting the target sentence through a feature extraction modeling-based algorithm, classify the parts of speech constituting the target sentence, extract vector values ​​for the characteristics of the sentence based on the classified parts of speech, and extract the intention for the target sentence by comparing the similarity with the vector values ​​of reference sentences classified by multiple intention categories.

[0057] In relation to the above, the feature extraction modeling may include modeling such as BoW and Word2Vec mentioned in the present invention.

[0058] In addition, the processor can analyze the target sentence through an algorithm based on embedding vector modeling included in the previously stored intent classification algorithm. At this time, the processor can learn a language model for multiple training sentences, extract an embedding vector value of the target sentence through the learned language model, and then extract the intent of the target sentence by comparing the similarity with the embedding vector value of a reference sentence (multiple training sentences after training is completed).

[0059] Finally, the processor can analyze the target sentence through a deep learning model included in the previously stored intent classification algorithm. At this time, the processor can extract the intent of the target sentence by performing inference on the target sentence after fine-tuning the language model for the plurality of training sentences with the addition of the target sentence, while the processor has trained the language model for the plurality of training sentences.

[0060] That is, the processor can identify the intent by analyzing the target sentence through one of the feature extraction modeling-based algorithm, the embedding vector modeling-based algorithm, and the deep learning modeling-based algorithm, and classify it into one of a plurality of intent categories.

[0061] According to one embodiment, when the algorithm learning completion step (step S103) is completed, the processor may perform a sentence violence judgment step (step S105).

[0062] In step S105, when the processor is in a state where the function of the algorithm learning completion step (step S103) is completed, if a target natural language to be determined for violence is input, the processor analyzes the target sentence corresponding to the target natural language based on the intent classification algorithm, calculates the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the plurality of intent categories, calculates a match value based on whether there is a match between the keywords constituting the target sentence and the keywords constituting the similar sentence, and compares the violence score of the target sentence calculated based on the similarity and the match value with a preset threshold value to determine whether the target sentence is violent.

[0063] According to one embodiment, when the target natural language is input from a user, the processor can analyze the target sentence through the intent classification algorithm.

[0064] In relation to the above, the processor can identify multiple morphemes included in the training sentence by analyzing the target sentence through natural language processing modeling included in the intent classification algorithm. Subsequently, the processor can identify the word segments constituting the target sentence based on the identified morphemes.

[0065] More specifically, the processor can identify multiple morphemes constituting the target sentence and distinguish which morpheme each morpheme is.

[0066] The types of the above-mentioned morphemes are classified into independent morphemes (morphemes that can be used alone (e.g., weather)), dependent morphemes (morphemes that depend on other words (e.g., ~eul, ~neun, ~da)), substantive morphemes (morphemes that have a substantive meaning (e.g., today)), and formal morphemes (morphemes that add grammatical relationships or formal meanings (e.g., particles, endings, affixes)), and the processor can analyze each type of the decomposed morphemes. At this time, the processor can distinguish and verify each type of the decomposed morphemes based on previously stored morpheme information.

[0067] According to one embodiment, when the processor completes the identification of a plurality of morphemes for the target sentence, it can identify a plurality of keywords constituting the target sentence based on the identified plurality of morphemes.

[0068] According to one embodiment, the processor may perform a tokenization process to identify at least one keyword within a training sentence. At this time, when performing the tokenization process, the processor may perform a morpheme tokenization method rather than word tokenization because, generally, unlike English, Korean is an agglutinative language in which morphemes are not composed solely of independent words.

[0069] According to one embodiment, the processor recognizes a plurality of morphemes and types of morphemes included in the target sentence, distinguishes the types of morphemes, recognizes a combination of independent morphemes and dependent morphemes as a single token, and can designate it as a single keyword.

[0070] According to one embodiment, the processor can identify multiple keywords included in the target sentence by performing a tokenization process of the morphological tokenization method. For example, the processor can identify words in "stop distributing work excessively to A." The processor can identify five keywords by performing morphological tokenization on "stop distributing work excessively to A": "to A" v "work" v "excessively" v "distributing" v "stop".

[0071] According to one embodiment, when the processor analyzes the target sentence through the intent classification algorithm and calculates a first vector value based on the meaning of the target sentence, it can identify an intent category (similar intent category) that includes a similar sentence, which is a reference sentence with the highest similarity to the target sentence among the plurality of intent categories, by performing a process of calculating similarity between the plurality of intent categories through the calculated first vector value.

[0072] That is, the processor can also calculate the similarity between the similar sentence and the target sentence through the similarity calculation process. At this time, the similarity calculated may be a value between 0 and 1, and may be a value calculated according to the degree of semantic similarity.

[0073] Afterward, when the calculation of similarity between the target sentence and the similar sentence is completed, the processor may calculate a first violence score by multiplying the calculated similarity by a preset first weight. In this regard, the first weight has a minimum value of 0 and a maximum value of 0.5, and may be fixed at 0.5 if there is no user setting.

[0074] According to one embodiment, when the calculation of the first violence score is completed, the processor can identify multiple keywords constituting the target sentence (e.g., first multiple keywords) and multiple keywords constituting the similar sentence (e.g., second multiple keywords) through natural language processing modeling of the intent classification algorithm.

[0075] In relation to the above, the processor can calculate a match value based on whether there is a match between the first plurality of keywords and the second plurality of keywords. More specifically, the processor can classify the matching keywords of the target sentence for the similar sentence as "1" and the non-matching keywords as "0" based on a boolean technique.

[0076] For example, the processor can calculate a match value of 0.2 when there are 5 keywords in the similar sentence and 4 keywords in the target sentence, and there are 2 matching keywords in the target sentence for the keywords in the similar sentence.

[0077] According to one embodiment, when the calculation of the match value is completed, the processor may calculate a second violence score by multiplying the calculated match value by a preset second weight. In this regard, the second weight has a minimum value of 0 and a maximum value of 0.5, and may be fixed at 0.5 if there is no user setting.

[0078] According to one embodiment, when the calculation of the first violence score and the second violence score is completed, the processor may calculate a sum score by summing the first violence score and the second violence score. (Sum score = First violence score (first weight (α) x similarity) + Second violence score (second weight (β) x Boolean match value) (wherein, the first weight (α) + second weight (β) is fixed at 1))

[0079] Afterwards, the processor compares the sum score with a preset threshold value matched to the similar intent category, and if the sum score exceeds the preset threshold value, it can determine that the target sentence is a sentence containing violence.

[0080] In addition, the processor can compare the sum score with a preset threshold value matched to the similar intent category, and if the sum score does not exceed the preset threshold value, determine that the target sentence is a sentence that does not contain violence.

[0081] FIG. 2 is a flowchart illustrating the vector information generation step of a learning and processing method for determining whether a sentence is violent according to an embodiment of the present invention.

[0082] Referring to FIG. 2, a learning and processing method for determining whether a sentence is violent, implemented as a computing device comprising one or more processors and one or more memories that store instructions that can be executed on said processors, may include a vector information generation step (e.g., the vector information generation step (step S101) of FIG. 1).

[0083] According to one embodiment, the vector information generation step may be a step of generating multiple vector information by analyzing each of the learning sentences corresponding to the multiple natural languages ​​through a previously stored vector analysis algorithm when multiple natural languages ​​are input from a user, and calculating a vector value for each of the learning sentences through the analysis result.

[0084] According to one embodiment, the vector information generation step may include a sentence classification step (step S201) and a vector calculation step (step S203) as detailed steps for performing the above-described function.

[0085] In step S201, the one or more processors (hereinafter referred to as processors) can classify the multiple natural words into the multiple intention categories through inputs separately received from the user while the multiple natural words are input from the user.

[0086] According to one embodiment, the processor may receive a separate input from the user while the plurality of natural words are input from the user. In this regard, the input received separately from the user may be an input that classifies the plurality of natural words into a plurality of intent categories.

[0087] For example, a first input separately entered by a user may be an input that classifies the first natural language among the plurality of natural languages, "Do you want to get hit?", into the first negative intent category among the plurality of intent categories to which the assault attribute is matched. As another example, a second input separately entered by a user may be an input that classifies the second natural language among the plurality of natural languages, "Give him all the bad work," into the second negative intent category among the plurality of intent categories to which the workplace harassment attribute (work attribute) is matched.

[0088] According to one embodiment, when the processor completes classifying the plurality of natural words by the plurality of intent categories through input separately received from the user, the vector calculation step (step S203) may be included.

[0089] In step S203, when the sentence classification step (step S201) is completed, the processor can analyze the training sentences corresponding to the plurality of natural languages ​​through the previously stored vector analysis algorithm, calculate an initial vector value based on the meaning of each of the plurality of training sentences, and generate vector information for each of the plurality of training sentences.

[0090] According to one embodiment, the processor can calculate an initial vector value corresponding to a training sentence by analyzing a training sentence corresponding to each of the plurality of natural languages ​​through a previously stored vector analysis algorithm. In this regard, the initial vector value may be a vector value calculated based on the meaning of the training sentence.

[0091] Accordingly, the processor can generate vector information (initial vector information) corresponding to each of the plurality of training sentences through the calculated initial vector value.

[0092] FIG. 3 is a block diagram illustrating the algorithm learning completion section of a learning and processing system for determining whether a sentence is violent according to an embodiment of the present invention.

[0093] Referring to FIG. 3, a learning and processing system for determining whether a sentence is violent, implemented as a computing device comprising one or more processors and one or more memories for storing instructions that can be executed by said processors, may include an algorithm learning completion unit (e.g., performing the same function as the algorithm learning completion step (step S103) of FIG. 1).

[0094] According to one embodiment, the algorithm learning completion unit (300) may be in the step of completing the learning of the intention classification algorithm by performing the grouping process for each of the identified vector values ​​after the generation of a plurality of vector information is completed, and then performing the grouping process for each of the plurality of vector information.

[0095] According to one embodiment, the algorithm learning completion unit (300) may include a grouping completion unit (301) and a learning completion unit (303) as detailed configurations for performing the above-described function.

[0096] According to one embodiment, when the generation of vector information for each of the plurality of learning sentences is completed, the grouping completion unit (301) can complete the grouping of the first learning sentence for the first intention category by correcting the first initial vector value of the first learning sentence classified as the first intention category among the plurality of intention categories to a central intention vector value derived based on the vector value of the reference sentence included in the first intention category.

[0097] According to one embodiment, the grouping completion unit (301) can start a grouping process for multiple learning sentences (301a) included in each of the multiple intention categories when the initial vector value for each of the multiple learning sentences is calculated by the performance of the function of the vector calculation unit (not shown) (e.g., performing the same function as the vector calculation step (step S203) of FIG. 2) and the generation of vector information (initial vector information) is completed.

[0098] More specifically, the grouping completion unit (301) can perform a grouping process for each of the plurality of intention categories when the function of the vector calculation unit is completed. In this regard, when the function of the vector calculation unit is completed, the grouping completion unit (301) can identify the first initial vector value of a first learning sentence (e.g., "give bad work to him") classified as a first intention category (e.g., first negative intention category) among the plurality of intention categories through the vector information.

[0099] According to one embodiment, when the identification of the first initial vector value for the first learning sentences classified into the first intention category is completed, the grouping completion unit (301) can correct the center intention vector value matched to the first intention category to the identified first initial vector value.

[0100] In relation to the above, the central intention vector value may be configured to be calculated through a vector value corresponding to each reference sentence among the reference sentences included for each of the plurality of intention categories, wherein the usage frequency is learned to be greater than or equal to a specified frequency.

[0101] For example, the reference sentences included in the first positive intention category (anti-bullying attribute category) may be various sentences containing meanings intended to prevent bullying. The curve reflected in the normal distribution table for the vector values ​​of the reference sentences included in the first positive intention category may be formed around reference sentences that are used at a frequency exceeding a specified level. That is, reference sentences based on vector values ​​corresponding to lines located far from the center curve are sentences with low usage frequency, and may be sentences with ambiguous meanings or those that are not generally used.

[0102] Accordingly, the above-mentioned central intention vector value may be a configuration calculated through the vector value for reference sentences (the central curve among the curves reflected in the normal distribution table) that are used with a frequency greater than a specified frequency among various reference sentences containing meanings for preventing bullying.

[0103] More specifically, the central intention vector value may be configured to be corrected to the initial vector value of the training sentence in order to prevent cases where the reference sentences (multiple training sentences used to train the intention classification algorithm in the vector information generation step and the algorithm learning completion step) which serve as a standard for determining whether the target sentence corresponding to the target natural language input in the sentence violence judgment step (e.g., sentence violence judgment step of FIG. 1 (step S105)) are not distinguished by the multiple intention categories and to set a standard for distinguishing by the multiple intention categories so that no omitted reference sentences occur.

[0104] That is, the above central intention vector value may be configured such that, as the learning of the above intention classification algorithm (303a) is completed, the previously classified learning sentences are set as reference sentences for each of the multiple intention categories, and the vector value corresponding to each reference sentence among the set reference sentences is calculated through a vector value that is learned with a usage frequency greater than or equal to a specified frequency.

[0105] According to one embodiment, the grouping completion unit (301) can correct the center intention vector value to the first initial vector value, and since the violence of the target sentence input thereafter is determined based on the vector value of the target sentence, accurate semantic interpretation of the learning sentence must be performed at the algorithm learning completion stage.

[0106] Accordingly, the grouping completion unit (301) may include training sentences that have initial vector values ​​of training sentences classified into the first intention category, which may have initial vector values ​​similar to other intention categories, and may refer to reference sentences that have an absolutely low frequency of use in the first intention category. When training the intention classification algorithm based on such reference sentences, it may hinder the determination of whether a target sentence input thereafter is violent.

[0107] Accordingly, the grouping completion unit (301) can complete a grouping process to more clearly distinguish the learning sentences classified into the first intention category from other intention categories by correcting the center intention vector value matched to the first intention category to the first initial vector value corresponding to the first learning sentences classified into the first intention category, thereby clearly distinguishing the learning sentences classified into the first intention category as sentences corresponding to the violence attribute matched to the first intention category and establishing a standard so that they are not confused with other categories.

[0108] According to one embodiment, the learning completion unit (303) can complete the learning of the intent classification algorithm (303a) when the grouping process of multiple learning sentences for each of the multiple intent categories is completed as the function of the grouping completion unit (301) is repeated multiple times.

[0109] According to one embodiment, as the function of the grouping completion unit (301) is repeated multiple times, the grouping process for each of the multiple intention categories is performed multiple times, and when the grouping process is completed, the learning of the intention classification algorithm (303a) can be completed.

[0110] In relation to the above, the intention classification algorithm (303a) may be an algorithm that analyzes the vector value of a target sentence to be input thereafter, classifies it into one of a plurality of intention classification categories, and calculates the vector value for the target sentence classified into one of the plurality of intention classification categories and the vector value for a plurality of keywords constituting the target sentence.

[0111] FIG. 4 is a flowchart illustrating the sentence violence determination step of a learning and processing method for determining whether a sentence is violent according to an embodiment of the present invention.

[0112] Referring to FIG. 4, a learning and processing method for determining whether a sentence is violent, implemented as a computing device comprising one or more processors and one or more memories that store instructions that can be executed by said processors, may include a sentence violence determination step (e.g., the sentence violence determination step (step S105) of FIG. 1).

[0113] According to one embodiment, the sentence violence determination step may be a step in which, when a target natural language to be determined for violence is input while the algorithm learning completion step (e.g., the algorithm learning completion step of FIG. 1 (step S103)) is completed, the target sentence corresponding to the target natural language is analyzed based on the intent classification algorithm to calculate the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the plurality of intent categories, and then calculates a match value based on whether there is a match between the keyword constituting the target sentence and the keyword constituting the similar sentence, and compares the violence score of the target sentence calculated based on the similarity and the match value with a preset threshold value to determine whether the target sentence is violent.

[0114] According to one embodiment, the sentence violence determination step may include a first vector value calculation step (step S401), a similarity intention category identification step (step S403), and a first violence score calculation step (step S405) as detailed steps for performing the above-described function.

[0115] In step S401, when the algorithm learning completion step is completed, the one or more processors (hereinafter referred to as processors) can analyze the target sentence corresponding to the target natural language through the vector analysis model of the intent classification algorithm and calculate a first vector value based on the meaning of the target sentence.

[0116] According to one embodiment, when the algorithm learning completion step is completed and the learning of the intent classification algorithm is completed, the processor can analyze a target sentence corresponding to the target natural language through the vector analysis model of the intent classification algorithm when a target natural language is input from a user.

[0117] Accordingly, the processor can calculate a first vector value based on the meaning of the target sentence based on the result of analyzing the target sentence.

[0118] According to one embodiment, when the calculation of the first vector value is completed, the processor may perform a similar intention category identification step (step S403).

[0119] In step S403, when the calculation of the first vector value is completed, the processor performs a process of calculating the similarity between the calculated first vector value and the vector value of a reference sentence included in the plurality of intention categories, thereby identifying a similar intention category that includes a similar sentence, which is the reference sentence with the highest similarity to the first vector value among the plurality of intention categories.

[0120] According to one embodiment, when the calculation of the first vector value is completed, the processor may perform a process of calculating similarity between the first vector value and the vector value of each reference sentence (composed of a learned sentence in the vector information generation step and the algorithm learning completion step) classified by the plurality of intent categories as the learning of the intent classification algorithm is completed.

[0121] According to one embodiment, the similarity process may be an algorithm in which the processor calculates the similarity between the first vector value and the vector value of each reference sentence classified by the plurality of intent categories through a stored similarity calculation algorithm.

[0122] In relation to the above, the similarity calculation algorithm may include at least one of a cosine similarity model, a Euclidean distance model, a Jaccard similarity model, and a Manhattan similarity model.

[0123] Accordingly, the processor can identify a reference sentence having the highest similarity with the first vector value by performing a similarity calculation process between the first vector value and the vector value of a reference sentence included in the plurality of intention categories. At this time, when the processor performs the similarity calculation process, the similarity between the first vector value and the vector value of a reference sentence included in the plurality of intention categories can be calculated as a similarity between 0 and 1.

[0124] For example, the processor can identify that the similarity between the first vector value and the first reference vector value, which is the vector value with the highest similarity to the first vector value, is 0.8, and then identify the first negative intention category (category of workplace bullying incitement attribute), which is an intention category containing a similar sentence having the first reference vector value, as a similar intention category.

[0125] In relation to the above, the similar intent category may be an intent category that includes reference sentences with a high degree of similarity to the meaning of the target sentence.

[0126] According to one embodiment, when the similar intent category identification step (step S403) is completed, the processor may perform the first violence score calculation step (step S405).

[0127] In step S405, the processor can calculate a first violence score, which is a configuration used to determine whether the target sentence is violent, based on the similarity calculated according to the function of the similar intention category identification step (e.g., the similar intention category identification step of FIG. 4 (step S403)) and a pre-set first weight.

[0128] According to one embodiment, when the identification of the similar intent category is completed, the processor can calculate a first violence score, which is the value obtained by multiplying the similarity identified according to the function of the similar intent category identification step (step S403)) and the first weight set in advance.

[0129] In relation to the above, the first weight set is configured to be matched to the similar intention category, and different weight values ​​may be matched to each violence attribute of the similar intention category (e.g., attribute encouraging workplace bullying (negative intention category), attribute preventing workplace bullying (positive intention category).

[0130] More specifically, when the processor completes the identification of a similar sentence, which is a reference sentence having the highest similarity to the vector value of a target sentence among reference sentences included in a similar intention category, it can confirm that the similarity of the first vector value to the first reference vector value is 0.8. In relation to the above, the similarity may be a configuration in which the similarity is calculated as a value between 0 and 1, and may be a configuration in which the similarity is calculated according to the degree of semantic similarity.

[0131] Subsequently, the processor can calculate a first violence score (first weight x similarity), which is a configuration used to calculate whether the target sentence is violent, based on the confirmed similarity and a pre-set first weight.

[0132] In relation to the above, the first violence score is a semantic score of the target sentence based on the meaning of the target sentence, and may be configured to be used to calculate a sum score to be compared with a pre-set threshold value that serves as a criterion for determining whether the target sentence is violent.

[0133] FIG. 5 is a flowchart illustrating the sentence violence determination step of a learning and processing method for determining whether a sentence is violent according to an embodiment of the present invention.

[0134] Referring to FIG. 5, a learning and processing method for determining whether a sentence is violent, implemented as a computing device comprising one or more processors and one or more memories that store instructions that can be executed by said processors, may include a sentence violence determination step (e.g., the sentence violence determination step (step S105) of FIG. 1).

[0135] According to one embodiment, the sentence violence determination step may be a step in which, when a target natural language to be determined for violence is input while the algorithm learning completion step (e.g., the algorithm learning completion step of FIG. 1 (step S103)) is completed, the target sentence corresponding to the target natural language is analyzed based on the intent classification algorithm to calculate the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the plurality of intent categories, and then calculates a match value based on whether there is a match between the keyword constituting the target sentence and the keyword constituting the similar sentence, and compares the violence score of the target sentence calculated based on the similarity and the match value with a preset threshold value to determine whether the target sentence is violent.

[0136] According to one embodiment, the sentence violence determination step may include a matching value calculation step (step S501), a second violence score calculation step (step S503), and a violence information provision step (step S505) as detailed steps for performing the above-described function.

[0137] In step S501, when the identification of the similar intent category is completed, the one or more processors (hereinafter referred to as processors) can perform a process of determining whether the keywords constituting the similar sentence included in the similar intent category and the keywords constituting the target sentence are identical, and can calculate a match value calculated according to whether each keyword matches.

[0138] According to one embodiment, when the identification of a similar intent category is completed by performing the function of a similar intent category identification step (e.g., the similar intent category identification step (step S403) of FIG. 4), the processor may perform a process of determining whether the keyword constituting the similar sentence included in the similar intent category is the same as the keyword constituting the target sentence.

[0139] According to one embodiment, when the identification of similar intent categories is completed, the processor can identify a similar sentence having a first reference vector value that has the highest similarity to the first vector value of the target sentence among the reference sentences included in the identified similar intent categories.

[0140] In relation to the above, the processor can identify multiple keywords constituting the target sentence and multiple keywords constituting the similar sentence by analyzing each of the target sentence and the similar sentence through a natural language processing model of an intent classification algorithm or a pre-stored natural language processing algorithm.

[0141] Afterward, when the processor completes the identification of multiple keywords constituting the target sentence and multiple keywords constituting the similar sentence, it can perform an identity verification process to check whether the multiple keywords of the target sentence and the multiple keywords of the similar sentence are identical.

[0142] For example, if the first keyword among the multiple keywords of the target sentence is the same as the first keyword among the multiple keywords of the similar sentence, a value "1" based on a Boolean technique for the first keyword can be calculated, and if the second keyword among the multiple keywords of the target sentence is not the same as the second keyword among the multiple keywords of the similar sentence, a value "0" based on a Boolean technique for the second keyword can be calculated.

[0143] At this time, when there are 4 keywords constituting the target sentence and 5 keywords constituting the similar sentence, if there are 2 keywords in the target sentence that match the keywords of the similar sentence (number of keywords with a value of "1" based on the Boolean method), a match value of 0.2 can be calculated.

[0144] Accordingly, when the processor completes the calculation of the matching value based on whether the multiple keywords of the target sentence and the multiple keywords of the similar sentence are identical, it can perform the second violence score calculation step (step S503).

[0145] In step S503, when the calculation of the match value is completed, the processor can calculate a second violence score, which is a configuration used to determine whether the target sentence is violent, based on the calculated match value and a pre-set second weight.

[0146] According to one embodiment, when the calculation of the match value is completed, the processor can calculate a second violence score, which is a configuration used to determine whether the target sentence is violent, based on the calculated match value and a preset second weight.

[0147] In relation to the above, the second violence score is a keyword score based on a keyword having violence among a plurality of keywords constituting the target sentence, and may have a different composition from the first violence score, which quantifies whether the target sentence has violence based on its meaning.

[0148] In relation to the above, the previously set second weight is configured to be matched to the similar intent category, and may have different weight values ​​matched to each violence attribute of the similar intent category (e.g., attribute encouraging workplace bullying (negative intent category), attribute preventing workplace bullying (positive intent category), and may be configured differently from the first weight reflected in the first violence weight based on the semantic value of the sentence. That is, the second weight may be configured to be reflected in the second violence weight based on the semantic value of the keyword constituting the sentence.

[0149] According to one embodiment, when the calculation of the first violence score and the second violence score is completed, the processor may perform a violence information provision step (step S505).

[0150] In step S505, when the calculation of the first violence score and the second violence score is completed, the processor calculates a sum score by summing the first violence score and the second violence score, checks whether the calculated sum score exceeds a preset threshold value matched to the similar intent category, and if it is confirmed through the check result that the sum score exceeds the preset threshold value, determines that violence exists in the target sentence, and can generate sentence violence information based on the history of the function execution of the intent classification algorithm and provide it to the user.

[0151] According to one embodiment, if the sum of the first violence score and the second violence score is 0.8 points and the preset threshold value matched to the similar intent category is 0.7 points, the processor can determine that the target sentence is a sentence having violence by confirming that the sum score exceeds the preset threshold value.

[0152] That is, the processor can determine that the target sentence is a target sentence having the violent nature of the attribute of inciting workplace bullying, which is an attribute of the first negative intention category.

[0153] According to one embodiment, when the processor determines that the target sentence is a sentence possessing violence, it can generate sentence violence information based on the history of the intent classification algorithm analyzing the target sentence.

[0154] In relation to the above, the sentence violence information is information generated through a history based on the performance of the function of the intent classification algorithm, and may be composed of the target sentence, whether the target sentence is violent, the intent category to which the target sentence is included, the reference sentence with the highest similarity to the target sentence among the reference sentences included in the intent category, and keywords containing a violent meaning among the keywords included in the target sentence.

[0155] For example, if the processor determines that the target sentence is a target sentence possessing the violence attribute of the workplace bullying incitement attribute of the first negative intent category, it may generate sentence violence information including the target sentence "Give bad work to him / her", the violence status of the target sentence as "true (existence)", the category as "first negative intent category (workplace bullying incitement attribute)", a similar sentence as "Let's give work to him / her / her and make him / her work overtime", and violence keywords as "bad work" and "give it to him / her".

[0156] Accordingly, the processor can provide the generated sentence violence information to the user to identify whether the target sentence is violent.

[0157] FIG. 6 is a drawing for explaining an example of the internal configuration of a computing device according to an embodiment of the present invention.

[0158] FIG. 6 illustrates an example of the internal configuration of a computing device according to an embodiment of the present invention. In the following description, descriptions of unnecessary embodiments that overlap with the descriptions of FIG. 1 to 5 described above will be omitted.

[0159] As illustrated in FIG. 6, the computing device (10000) may include at least one processor (11100), memory (11200), peripheral interface (11300), input / output subsystem (11400), power circuit (11500), and communication circuit (11600). In this case, the computing device (10000) may correspond to a user terminal (A) connected to a haptic interface device or the aforementioned computing device (B).

[0160] The memory (11200) may include, for example, high-speed random access memory, a magnetic disk, SRAM, DRAM, ROM, flash memory, or non-volatile memory. The memory (11200) may include software modules, instruction sets, or various other data required for the operation of the computing device (10000).

[0161] At this time, access to memory (11200) from other components, such as the processor (11100) or peripheral device interface (11300), can be controlled by the processor (11100).

[0162] The peripheral device interface (11300) can connect input and / or output peripheral devices of the computing device (10000) to the processor (11100) and memory (11200). The processor (11100) can perform various functions for the computing device (10000) and process data by executing software modules or instruction sets stored in the memory (11200).

[0163] The input / output subsystem (11400) can connect various input / output peripherals to the peripheral interface (11300). For example, the input / output subsystem (11400) may include a controller for connecting peripherals such as a monitor, keyboard, mouse, printer, or, if necessary, a touchscreen or sensor to the peripheral interface (11300). According to another aspect, input / output peripherals may be connected to the peripheral interface (11300) without passing through the input / output subsystem (11400).

[0164] The power circuit (11500) can supply power to all or part of the components of the terminal. For example, the power circuit (11500) may include one or more power sources such as a power management system, a battery or alternating current (AC), a charging system, a power failure detection circuit, a power converter or inverter, a power status indicator, or any other components for power generation, management, and distribution.

[0165] The communication circuit (11600) can enable communication with another computing device using at least one external port.

[0166] Alternatively, as described above, the communication circuit (11600) may enable communication with other computing devices by including an RF circuit and transmitting and receiving an RF signal, also known as an electromagnetic signal.

[0167] The embodiment of FIG. 6 is merely an example of a computing device (10000), and the computing device (11000) may have some components shown in FIG. 6 omitted, additional components not shown in FIG. 6 added, or a configuration or arrangement that combines two or more components. For example, a computing device for a communication terminal in a mobile environment may include, in addition to the components shown in FIG. 6, a touchscreen or a sensor, etc., and the communication circuit (1160) may include a circuit for RF communication of various communication methods (WiFi, 3G, LTE, Bluetooth, NFC, Zigbee, etc.). The components that can be included in the computing device (10000) may be implemented as hardware, software, or a combination of both hardware and software, including one or more integrated circuits specialized for signal processing or applications.

[0168] Methods according to embodiments of the present invention may be implemented in the form of program instructions that can be executed through various computing devices and recorded on a computer-readable medium. In particular, the program according to the present embodiment may be configured as a PC-based program or an application dedicated to a mobile terminal. An application to which the present invention is applied may be installed on a user terminal through a file provided by a file distribution system. For example, the file distribution system may include a file transmission unit (not shown) that transmits the file upon a request from the user terminal.

[0169] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0170] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed across networked computing devices and stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0171] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0172] Although the embodiments have been described above with reference to limited embodiments and drawings, those skilled in the art can make various modifications and variations from the description above. For example, appropriate results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents. Therefore, other implementations, other embodiments, and equivalents to the claims below also fall within the scope of the claims.

Claims

Claim 1 A learning and processing method for determining the violent nature of a sentence, implemented by a computing device comprising one or more processors and one or more memories for storing instructions executable by said processors, comprising: a vector information generation step in which, when multiple natural languages ​​are input from a user, each learning sentence corresponding to said multiple natural languages ​​is analyzed through a previously stored vector analysis algorithm, and a vector value for each of said learning sentences is calculated through said analysis results to generate multiple vector information; and an algorithm learning completion step in which, when the generation of the multiple vector information is completed, each vector value of said multiple vector information is identified, and a grouping process is performed for each of said identified vector values, wherein the grouping process is performed according to multiple intent categories to complete the learning of an intent classification algorithm. A sentence violence determination step, wherein, when a target natural language for determining violence is input while the algorithm learning completion step is completed, based on the intent classification algorithm, a target sentence corresponding to the target natural language is analyzed to calculate the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the multiple intent categories, and then a match value is calculated based on whether there is a match between the keyword constituting the target sentence and the keyword constituting the similar sentence, and the violence score of the target sentence calculated based on the similarity and the match value is compared with a preset threshold value to determine whether the target sentence is violent; wherein the vector information generation step includes a sentence classification step, wherein, while multiple natural languages ​​are input from a user, the multiple natural languages ​​are classified according to the multiple intent categories through input separately received from the user;The method comprises: a vector calculation step in which, when the sentence classification step is completed, the learning sentences corresponding to the plurality of natural languages ​​are analyzed through the previously stored vector analysis algorithm to calculate an initial vector value based on the meaning of each of the plurality of learning sentences and generate vector information for each of the plurality of learning sentences; wherein the algorithm learning completion step comprises: a grouping completion step in which, when the generation of vector information for each of the plurality of learning sentences is completed, the first initial vector value of the first learning sentence classified into the first intention category among the plurality of intention categories is corrected to a central intention vector value derived based on the vector value of a reference sentence included in the first intention category, thereby completing the grouping of the first learning sentence for the first intention category; and a learning completion step in which, when the grouping process of the plurality of learning sentences for each of the plurality of intention categories is completed as the function of the grouping completion step is repeated multiple times, the learning of the intention classification algorithm is completed....including, wherein the central intention vector value is configured to be calculated through a vector value corresponding to each reference sentence learned with a usage frequency exceeding a specified frequency among the reference sentences included for each of the plurality of intention categories, and is configured to be corrected to the initial vector value of the training sentence to set criteria for distinguishing by the plurality of intention categories so as not to occur, and to prevent cases where reference sentences serving as criteria for determining the violent nature of a target sentence corresponding to the target natural language input in the sentence violence judgment step are not distinguished by the plurality of intention categories, and so as not to occur where reference sentences are omitted, wherein the plurality of intention categories include a positive intention category and a negative intention category distinguished based on whether the reference sentence has violent attributes, and is configured such that the central intention vector value is matched to each of the positive intention category and the negative intention category, and includes a reference sentence having a vector value for deriving the central intention vector value, wherein when the training of the intention classification algorithm is completed according to the performance of the function of the algorithm learning completion step, the training sentences corresponding to the initial vector value are classified as reference sentences and included in the positive intention category and the negative intention category, and the sentence violence judgment step is... A first vector value calculation step in which, when the algorithm learning completion step is completed and the target natural language is input, a target sentence corresponding to the target natural language is analyzed through the intent classification algorithm to calculate a first vector value based on the meaning of the target sentence;When the calculation of the first vector value is completed, a similarity calculation process is performed between the calculated first vector value and the vector value of a reference sentence included in the plurality of intention categories, and a similar intention category identification step is performed to identify a similar sentence, which is a reference sentence having the highest vector value with respect to the first vector value among the reference sentences included in each of the plurality of intention categories, along with the intention category containing the similar sentence, as a similar intention category; The method includes a first violence score calculation step, which calculates a first violence score, which is a component used to determine whether the target sentence is violent, based on a similarity calculated according to the function of the similar intent category identification step and a preset first weight; wherein the sentence violence determination step comprises: a match value calculation step, which, when the identification of the similar intent category is completed, performs a process of determining identity between a keyword constituting a similar sentence included in the similar intent category and a keyword constituting the target sentence, and calculates a match value calculated according to whether each keyword matches; and a second violence score calculation step, which, when the calculation of the match value is completed, calculates a second violence score, which is a component used to determine whether the target sentence is violent, based on the calculated match value and a preset second weight. And when the calculation of the first violence score and the second violence score is completed, a sum score is calculated by summing the first violence score and the second violence score, and after checking whether the calculated sum score exceeds a preset threshold value matched to the similar intent category, if it is confirmed through the check result that the sum score exceeds the preset threshold value, it is determined that violence exists in the target sentence, and a violence information provision step is provided to the user by generating sentence violence information through a history based on the function execution of the intent classification algorithm;A learning and processing method for determining the violence of a sentence, further comprising, wherein the sentence violence information is information generated through a history based on the performance of the function of the intent classification algorithm, and is characterized by including the target sentence, whether the target sentence is violent, the intent category to which the target sentence is included, the reference sentence with the highest similarity to the target sentence among the reference sentences included in the intent category, and keywords containing a violent meaning among the keywords included in the target sentence. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 delete Claim 7 delete Claim 8 delete Claim 9 A learning and processing device for determining the violent nature of a sentence, implemented as a computing device comprising one or more processors and one or more memories for storing instructions executable by said processors, wherein when a plurality of natural languages ​​are input from a user, a vector information generation unit analyzes each of the learning sentences corresponding to said plurality of natural languages ​​through a previously stored vector analysis algorithm, calculates a vector value for each of said learning sentences through the analysis result, and generates a plurality of vector information; and an algorithm learning completion unit, wherein when the generation of the plurality of vector information is completed, identifies the vector value of each of said plurality of vector information and performs a grouping process for each of said identified vector values, and performs the grouping process according to a plurality of intent categories to complete the learning of an intent classification algorithm. A sentence violence determination unit that, when the function of the algorithm learning completion unit is completed and a target natural language for determining violence is input, analyzes a target sentence corresponding to the target natural language based on the intent classification algorithm, calculates the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the plurality of intent categories, calculates a match value based on whether there is a match between the keyword constituting the target sentence and the keyword constituting the similar sentence, and determines the violence of the target sentence by comparing the violence score of the target sentence calculated based on the similarity and the match value with a preset threshold value; wherein the vector information generation unit includes a sentence classification unit that classifies the plurality of natural languages ​​according to the plurality of intent categories through input separately received from the user when the plurality of natural languages ​​are input from the user;The method includes: a vector calculation unit that, when the sentence classification unit is completed, analyzes the learning sentences corresponding to the plurality of natural languages ​​through the previously stored vector analysis algorithm to calculate an initial vector value based on the meaning of each of the plurality of learning sentences and generates vector information for each of the plurality of learning sentences; wherein the algorithm learning completion unit comprises: a grouping completion unit that, when the generation of vector information for each of the plurality of learning sentences is completed, corrects the first initial vector value of the first learning sentence classified into the first intention category among the plurality of intention categories to a central intention vector value derived based on the vector value of a reference sentence included in the first intention category, thereby completing the grouping of the first learning sentence for the first intention category; and a learning completion unit that completes the learning of the intention classification algorithm when the grouping process of the plurality of learning sentences for each of the plurality of intention categories is completed as the function of the grouping completion unit is repeated multiple times....including, wherein the central intention vector value is configured to be calculated through a vector value corresponding to each reference sentence learned with a usage frequency exceeding a specified frequency among the reference sentences included for each of the plurality of intention categories, and is configured to be corrected to the initial vector value of the training sentence to set criteria for distinguishing by the plurality of intention categories so as not to occur, and to prevent cases where reference sentences serving as criteria for determining the violent nature of a target sentence corresponding to the target natural language input to the sentence violence judgment unit are not distinguished by the plurality of intention categories, and so as not to occur where reference sentences are omitted, wherein the plurality of intention categories include a positive intention category and a negative intention category distinguished based on whether the reference sentence has violent attributes, and is configured such that the central intention vector value is matched to each of the positive intention category and the negative intention category, and includes a reference sentence having a vector value for deriving the central intention vector value, wherein when the training of the intention classification algorithm is completed according to the function execution of the algorithm learning completion unit, the training sentences corresponding to the initial vector value are classified as reference sentences and included in the positive intention category and the negative intention category, and the sentence violence judgment unit A first vector value calculation unit that, when the algorithm learning completion unit is completed and the target natural language is input, analyzes a target sentence corresponding to the target natural language through the intent classification algorithm and calculates a first vector value based on the meaning of the target sentence;When the calculation of the first vector value is completed, a similarity category identification unit performs a similarity calculation process between the calculated first vector value and the vector value of a reference sentence included in the plurality of intention categories, and identifies the similar sentence, which is a reference sentence having the highest vector value with respect to the first vector value among the reference sentences included in each of the plurality of intention categories, along with the intention category containing the similar sentence, as a similar intention category; and a first violence score calculation unit that calculates a first violence score, which is a component used to determine whether the target sentence is violent, based on a similarity calculated according to the function of the similar intent category identification unit and a preset first weight; wherein the sentence violence determination unit comprises: a match value calculation unit that, when the identification of the similar intent category is completed, performs a process of determining identity between a keyword constituting a similar sentence included in the similar intent category and a keyword constituting the target sentence, and calculates a match value calculated according to whether each keyword matches; and a second violence score calculation unit that, when the calculation of the match value is completed, calculates a second violence score, which is a component used to determine whether the target sentence is violent, based on the calculated match value and a preset second weight. When the calculation of the first violence score and the second violence score is completed, a sum score is calculated by summing the first violence score and the second violence score, and after checking whether the calculated sum score exceeds a preset threshold value matched to the similar intent category, if it is confirmed through the check result that the sum score exceeds the preset threshold value, it is determined that violence exists in the target sentence, and a violence information providing unit generates sentence violence information through a history based on the performance of the intent classification algorithm function and provides it to the user;A learning and processing device for determining the violence of a sentence, characterized in that it further includes, wherein the sentence violence information is information generated through a history based on the performance of the function of the intent classification algorithm, and includes the target sentence, whether the target sentence is violent, the intent category to which the target sentence is included, the reference sentence with the highest similarity to the target sentence among the reference sentences included in the intent category, and keywords containing a violent meaning among the keywords included in the target sentence. Claim 10 As a computer-readable recording medium, the computer-readable recording medium stores instructions for a computing device to perform the following steps, the steps comprising: a vector information generation step in which, when multiple natural languages ​​are input by a user, each learning sentence corresponding to the multiple natural languages ​​is analyzed through a previously stored vector analysis algorithm, and a vector value for each learning sentence is calculated through the analysis result to generate multiple vector information; and an algorithm learning completion step in which, when the generation of the multiple vector information is completed, each vector value of the multiple vector information is identified, and a grouping process is performed for each of the identified vector values, wherein the grouping process is performed by multiple intent categories to complete the learning of the intent classification algorithm. A sentence violence determination step, wherein, when a target natural language for determining violence is input while the algorithm learning completion step is completed, based on the intent classification algorithm, a target sentence corresponding to the target natural language is analyzed to calculate the similarity of the target sentence to a similar sentence which is one of the reference sentences included in the similar intent category among the multiple intent categories, and then a match value is calculated based on whether there is a match between the keyword constituting the target sentence and the keyword constituting the similar sentence, and the violence score of the target sentence calculated based on the similarity and the match value is compared with a preset threshold value to determine whether the target sentence is violent; wherein the vector information generation step includes a sentence classification step, wherein, while multiple natural languages ​​are input from a user, the multiple natural languages ​​are classified according to the multiple intent categories through input separately received from the user;The method comprises: a vector calculation step in which, when the sentence classification step is completed, the learning sentences corresponding to the plurality of natural languages ​​are analyzed through the previously stored vector analysis algorithm to calculate an initial vector value based on the meaning of each of the plurality of learning sentences and generate vector information for each of the plurality of learning sentences; wherein the algorithm learning completion step comprises: a grouping completion step in which, when the generation of vector information for each of the plurality of learning sentences is completed, the first initial vector value of the first learning sentence classified into the first intention category among the plurality of intention categories is corrected to a central intention vector value derived based on the vector value of a reference sentence included in the first intention category, thereby completing the grouping of the first learning sentence for the first intention category; and a learning completion step in which, when the grouping process of the plurality of learning sentences for each of the plurality of intention categories is completed as the function of the grouping completion step is repeated multiple times, the learning of the intention classification algorithm is completed....including, wherein the central intention vector value is configured to be calculated through a vector value corresponding to each reference sentence learned with a usage frequency exceeding a specified frequency among the reference sentences included for each of the plurality of intention categories, and is configured to be corrected to the initial vector value of the training sentence to set criteria for distinguishing by the plurality of intention categories so as not to occur, and to prevent cases where reference sentences serving as criteria for determining the violent nature of a target sentence corresponding to the target natural language input in the sentence violence judgment step are not distinguished by the plurality of intention categories, and so as not to occur where reference sentences are omitted, wherein the plurality of intention categories include a positive intention category and a negative intention category distinguished based on whether the reference sentence has violent attributes, and is configured such that the central intention vector value is matched to each of the positive intention category and the negative intention category, and includes a reference sentence having a vector value for deriving the central intention vector value, wherein when the training of the intention classification algorithm is completed according to the performance of the function of the algorithm learning completion step, the training sentences corresponding to the initial vector value are classified as reference sentences and included in the positive intention category and the negative intention category, and the sentence violence judgment step is... A first vector value calculation step in which, when the algorithm learning completion step is completed and the target natural language is input, a target sentence corresponding to the target natural language is analyzed through the intent classification algorithm to calculate a first vector value based on the meaning of the target sentence;When the calculation of the first vector value is completed, a similarity calculation process is performed between the calculated first vector value and the vector value of a reference sentence included in the plurality of intention categories, and a similar intention category identification step is performed to identify a similar sentence, which is a reference sentence having the highest vector value with respect to the first vector value among the reference sentences included in each of the plurality of intention categories, along with the intention category containing the similar sentence, as a similar intention category; The method includes a first violence score calculation step, which calculates a first violence score, which is a component used to determine whether the target sentence is violent, based on a similarity calculated according to the function of the similar intent category identification step and a preset first weight; wherein the sentence violence determination step comprises: a match value calculation step, which, when the identification of the similar intent category is completed, performs a process of determining identity between a keyword constituting a similar sentence included in the similar intent category and a keyword constituting the target sentence, and calculates a match value calculated according to whether each keyword matches; and a second violence score calculation step, which, when the calculation of the match value is completed, calculates a second violence score, which is a component used to determine whether the target sentence is violent, based on the calculated match value and a preset second weight. And when the calculation of the first violence score and the second violence score is completed, a sum score is calculated by summing the first violence score and the second violence score, and after checking whether the calculated sum score exceeds a preset threshold value matched to the similar intent category, if it is confirmed through the check result that the sum score exceeds the preset threshold value, it is determined that violence exists in the target sentence, and a violence information provision step is provided to the user by generating sentence violence information through a history based on the function execution of the intent classification algorithm;A computer-readable recording medium further comprising, wherein the sentence violence information is information generated through a history based on the performance of the function of the intent classification algorithm, and is characterized by including the target sentence, whether the target sentence is violent, the intent category to which the target sentence is included, the reference sentence with the highest similarity to the target sentence among the reference sentences included in the intent category, and keywords containing a violent meaning among the keywords included in the target sentence.;