Harmful statement identification method for appointment software private chat function
By establishing an emoji list and using a sliding window mechanism to detect input messages in the private chat function of the Epic Software, the problem that the existing technology cannot effectively identify bad information in emojis is solved, and effective identification and blocking of bad information is achieved to ensure safe communication between users.
Patent Information
- Application Number
- CN202411905771.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-23
AI Technical Summary
The existing chat function of the app is unable to effectively detect the possible negative information in emoticons, causing people with bad intentions to use emoticons to bypass the detection system to transmit information about harassment, fraud, violence and other meanings.
By creating an emoji list, marking the extended meaning of each emoji and its usage frequency probability value, replacing the emoji in the input message with encoding, using the sliding window mechanism to detect the message, and blocking the message from sending and reminding the user if there are illegal words.
It realizes effective identification and blocking of bad information in emoticons, prevents people with bad intentions from using emoticons to transmit bad information, and ensures that users can communicate safely in private chat function.
Smart Images

Figure CN120068847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and specifically to a method for identifying harmful statements in the private chat function of a dating photography software. Background Art
[0002] Dating photography software usually has a built-in chat function for further communication between users. The chat function is an important tool for users to communicate about the details of dating photography. Since the users of dating photography software are mostly young people, including professional practitioners such as models, makeup artists, photographers, and other users with photography needs, in addition to using ordinary text expressions during chatting, emoji are often used to express emotions and information. Emoji has the characteristics of being concise and having multiple meanings, and its colors are bright and the design is memorable, which is deeply loved by young people and is often used in communication.
[0003] However, precisely because of the characteristics of emoji, some people with malicious intentions often use these emoji, or use a combination of emoji or a combination of emoji and text to convey information with meanings such as harassment, fraud, and violence. Currently, some software with chat functions also have a built-in harmful information detection system. However, these detection systems usually only target text and cannot detect the possible implicit harmful information in emoji. Some people will use this method to bypass the platform's detection system and convey harmful information.
[0004] At the same time, the existing detection is usually carried out when the send button is pressed. However, this method has problems. When it is prompted that the input content contains harmful content / sensitive content / content violating the community convention, the user may not know where it is sensitive and thus have no idea how to modify it.
[0005] Therefore, there is an urgent need for a method for identifying harmful statements in the private chat function of a dating photography software to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for identifying harmful statements in the private chat function of a dating photography software to solve the problems raised in the above background art.
[0007] To solve the above technical problems, the present invention provides the following technical solution. A method for identifying harmful statements in the private chat function of a dating photography software includes the following steps:
[0008] Step A: Establish an emoji list and mark the usage frequency probability value for each extended meaning of the same emoji;
[0009] Step B: Replace the emoji in the input message with the corresponding encoding;
[0010] Step C: Set up detection conditions. If the input message does not meet the conditions, supplement and combine the input message. If the input message meets the conditions, directly perform detection;
[0011] Step D: Use the sliding window mechanism to detect the input message after deformation. If there are any illegal words in the input message after inspection, remind the user and prevent the message from being sent.
[0012] Preferably, it further includes:
[0013] Step E: Update the list of high-risk emojis according to the usage feedback of customers, and regularly update the usage frequency of each extended meaning of the emojis.
[0014] Preferably, the emoji in step A has a unique corresponding code, and multiple extended meanings correspond to the code corresponding to the unique emoji.
[0015] Preferably, step C includes:
[0016] Step C1: Set up detection conditions not less than four bytes;
[0017] Step C2: If the input message does not meet the detection conditions, combine the already sent message with the input message until the length of the unpassed input message meets the detection conditions;
[0018] Step C3: Detect the input message that meets the conditions.
[0019] Preferably, step D includes:
[0020] Step D1: Set the sliding window to 6 coding units, and set the step size to one coding unit;
[0021] Step D2: When the length of the input message exceeds the capacity of the sliding window, the sliding window starts to slide and detect the input message. Starting from the starting position of the input message, each time it slides, it intercepts a string of the window size, and slides one step size each time;
[0022] Step D3: Each time the sliding window slides, a substring will be generated in the string. If there is an emoji in the substring, find the corresponding meaning group according to the emoji code, replace the emoji content with the meaning, and perform semantic detection on the message after replacement;
[0023] Step D4: If harmful information is detected in the statement, prevent the message from being sent and point out the illegal words in the input message to the customer.
[0024] Preferably, it further includes:
[0025] Step D5. Repeat steps D1 - D4. If no harmful information is detected, the message is allowed to be sent.
[0026] Preferably, the sliding window detection in step D is performed simultaneously with the user's input message, and the customer is reminded in real time.
[0027] Preferably, step D3 further includes:
[0028] When performing the replacement of meaning groups, the initial meaning of the emoji is preferably selected and replaced into the sentence first, and then the extended meaning is used for replacement and the semantics of the replaced message is detected.
[0029] Preferably, the replacement order of the extended meanings is determined according to the probability value of the usage frequency of the tags, and the extended meaning with a higher probability value is preferentially replaced.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] 1. The present invention replaces the emojis existing in the message with a uniquely corresponding code, thereby transforming the entire sentence into a whole string for processing; uses a sliding window mechanism to detect the transformed sentence, replaces the emoji codes inside the sliding window with multiple extended meanings, and performs semantic detection on each replaced extended meaning to prevent malicious people from using the combination of emojis or the combination of emojis and text to transmit information with harassing, fraudulent, and violent meanings.
[0032] 2. The present invention performs real-time detection on the input message through the sliding window mechanism, realizes detection while inputting, and the user can timely know the illegal words in the input message and can timely modify the illegal words.
[0033] 3. For the semantic replacement order of emojis, the present invention is determined according to the probability value of the usage frequency of the tags, and the extended meaning with a higher probability value is preferentially replaced, preventing malicious people from using different extended meanings of emojis to transmit information with harassing, fraudulent, and violent meanings.
[0034] 4. The present invention updates the high-risk list according to the user's feedback. At the same time, the usage frequency of each extended meaning of the emoji should regularly recapture data for calculation to update the usage probability of each extended meaning of the emoji, ensuring coverage of the latest usage trend of the extended meaning and preventing malicious people from using the new extended meaning of the emoji to transmit information with harassing, fraudulent, and violent meanings. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flowchart of the harmful statement recognition method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0036] To facilitate the solution of the problems described in the background art, an embodiment of the present invention provides a harmful statement recognition method for the private chat function of a dating software. Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0037] Please refer to Figure 1 , this embodiment provides a harmful statement recognition method for the private chat function of a dating software, including the following steps:
[0038] Step A: Establish a list of high-risk emojis for storing some emojis that are likely to be associated with bad content in a specific context (whether in the case of consecutive emojis or in combination with text content).
[0039] When an emoji is initially designed, the designer assigns it an initial meaning. However, in actual use, these emojis will gradually develop new meanings. The number of these meanings is usually limited and there are differences in the degree of common use. Therefore, in addition to the initial meaning, a usage frequency probability value is marked for each extended meaning of the same emoji. This usage frequency probability value is calculated by analyzing sentences using the corresponding emoji on the Internet (for example, comments posted on social media).
[0040] Step B: When processing emojis, they are not recognized in the form of images, but in the form of a string of special codes, such as Unicode encoding, rich text format encoding, or custom encoding, etc. Therefore, establish a mapping relationship between the emoji and its code, replace the emojis existing in the message in the input box with the corresponding codes, so as to transform the entire message in the input box into a whole string for processing;
[0041] Step C: Set detection conditions. If the input message does not meet the conditions, the input message is supplemented and combined. The input message that meets the conditions is directly detected;
[0042] Step D: Use the sliding window mechanism to detect the input message that has been transformed. If the input message contains illegal words after inspection, the user is reminded, and the user modifies the input message according to the reminder;
[0043] Specifically, it further includes:
[0044] Step E: Allow users to provide feedback to adjust and optimize the judgment rules for semantic detection. If a user believes that a message is risk-free but is considered to contain harmful information, or believes that a message is harmful but fails to be processed by the recognition system, they can provide feedback on this. The review staff will update the high-risk emoji list based on this information.
[0045] At the same time, data on the usage frequency of each extended meaning of emojis should be regularly recaptured and calculated to update the usage probability of the extended meaning, ensuring coverage of the latest usage trends of the extended meaning.
[0046] Specifically, the emojis in step A have a unique corresponding encoding, and the encoding corresponding to the unique emoji corresponds to multiple extended meanings.
[0047] Specifically, step C includes:
[0048] Step C1, set the detection condition: not less than four bytes (generally, emojis take up 4 bytes when encoded in UTF-8; English letters occupy 1 character, and Chinese characters occupy 3 characters; 4 characters is the maximum occupancy for a single input).
[0049] Step C2, if the input message does not meet the detection condition, the previously sent message will be combined with the input message until the length of the unpassed input message meets the detection condition, and it will be regarded as a single message for judgment to prevent the behavior of evading the harmful information recognition mechanism by splitting a complete sentence into pieces. The number of previous messages taken is judged according to the length of the previous message. If the length of the previous message is also too short, one more previous message will be taken upward until the length of a single message meets the requirement, or all messages are taken, or the maximum value of the number of messages preset in the program is reached.
[0050] Step C3, messages that meet the detection condition are directly detected.
[0051] Specifically, step D includes:
[0052] Step D1, set the sliding window to 6 coding units (a coding unit refers to the smallest unit that makes up a string in a message, such as a Chinese character, an emoji, two bytes, etc., which can be specifically set according to the user's needs), and at the same time set the step size to one coding unit, which can be specifically set according to the needs.
[0053] Step D2, when the input message exceeds the capacity of the sliding window, the sliding window starts to slide and detect the input message. Starting from the starting position of the input message, each time it slides, a string of the window size is intercepted, and each time it slides one step size.
[0054] When the number of characters entered by the user exceeds the capacity of the sliding window, the content in the sliding window starts to be detected and the window starts to slide. Starting from the starting position of the input content, each slide intercepts a string of the window size, with a step size of one character each time. This step is repeated until all parts of the entire input content are covered and detected by the sliding window. This method allows for detection to start when the user is entering content, enabling detection while inputting.
[0055] Step D3: Each time the sliding window slides, a substring is generated in the string. If there is an emoji in the substring, the corresponding meaning group is found according to the emoji encoding, and after replacing the emoji content with the meaning, semantic detection is performed on the message after the replacement.
[0056] Step D4: After semantic detection, only harmless content can be sent. If the statement is detected to contain harmful information, the message sending is blocked, and the illegal words in the input message are pointed out to the customer.
[0057] Specifically, it also includes:
[0058] Step D5: Repeat steps D1 - D4. If no harmful information is detected, the message is allowed to be sent.
[0059] Specifically, the sliding window detection in step D is performed simultaneously with the user's input message, and the customer is reminded in real time to modify the illegal words detected in the input message. The customer can know which words in the input message are illegal words and make timely modifications to the illegal words.
[0060] Specifically, step D3 further includes:
[0061] When performing the replacement of the meaning group, the initial meaning of the emoji is preferably selected and inserted into the sentence, and then the extended meaning is used for replacement and semantic detection is performed on the message after the replacement.
[0062] Specifically, the replacement order of the extended meanings is determined according to the marked usage frequency probability values. The extended meaning with a higher probability value is preferentially replaced.
[0063] During use:
[0064] Before detecting the input message, a high - risk emoji list is first established to store some emojis that are likely to be associated with bad content in a specific context. A usage frequency probability value is marked for each extended meaning of the same emoji, and this usage frequency probability value is calculated by analyzing sentences with the corresponding emoji on the network (for example, comments posted on social media).
[0065] While the user is inputting information using a social software, the emoticons in the message are replaced with their uniquely corresponding codes, thus transforming the entire input information into a single string for processing.
[0066] After completing the processing of the message, it is judged whether the length of the message in the input chat box meets the detection condition. If the input message meets the requirement of not less than four bytes, the next step is to directly check the processed message; if the message does not meet the requirement of not less than four bytes, the sent message will be combined with the input message until the length of the unqualified input message meets the detection condition, and it is regarded as a single message for judgment, so as to prevent the behavior of evading the harmful information recognition mechanism by splitting a complete sentence into pieces. The number of previous messages taken is judged according to the length of the previous message. If the length of the previous message is also too short, one more previous message will be taken upward until the length of a single message meets the requirement, or all the messages are taken, or the maximum value of the number of messages preset in the program is reached. Process the information that does not meet the conditions to make the message convenient for detection.
[0067] For the message that meets the detection condition, a sliding window mechanism is used for detection. Set the sliding window to 6 coding units, and at the same time set the step size of the sliding window to one coding unit. When the number of characters input by the user exceeds the capacity of the sliding window, start to detect the content in the sliding window and start sliding. Starting from the starting position of the input content, each time it slides, a string of the window size is intercepted, and each time it slides by one step size. This step is repeated until all parts of the entire input content are covered and detected by the sliding window. This method allows detection to start when the user is inputting content. Compared with the usual method of starting detection after pressing the send button, it can reduce the delay of the statement being sent out, which is suitable for real-time output scenarios such as chatting.
[0068] Each time the sliding window slides, a new substring will be generated in the string, and each substring is analyzed. If there is an emoticon in the substring, find its corresponding extended meaning according to its code, and replace the emoticon content with the extended meaning for judgment. When replacing, first choose its initial meaning and substitute it into the sentence; then use the extended meaning for replacement and perform semantic detection. The replacement order of the extended meanings is determined according to the previously marked frequency probability values of the extended meanings, and the extended meaning with a higher probability value is preferentially replaced. Perform semantic detection on the message for the replaced extended meaning to ensure that any extended meaning of the emoticon used in the message will not generate meanings such as harassment, fraud, and violence.
[0069] After the detection of the message, only harmless messages can be sent. If harmful information is detected in the statement, the message sending will be blocked, and the illegal words in the input message will be pointed out to the customer. In this solution, real-time detection can be achieved during input, so that users can timely know where in their messages are intercepted by the monitoring system and correct the content input into the chat box for the illegal words.
[0070] During the use of this method, users are allowed to provide feedback to adjust and optimize the judgment rules of semantic detection. If a user believes that a message is risk-free but is considered to contain harmful information, or believes that the information is harmful but has not been processed by the recognition system, they can provide feedback on it, and the reviewers will update the high-risk list based on this information. At the same time, the usage frequencies of the various extended meanings of emoji should also be recalculated by regularly scraping data to update the usage probabilities of the extended meanings and ensure coverage of the latest usage trends of the extended meanings.
[0071] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for identifying harmful sentences in the private chat function of a dating software, characterized in that: The steps include: Step A: Create an emoji list and mark the usage frequency probability value for each extended meaning of the same emoji; Step B: Replace the emoticons in the input message with corresponding codes; Step C: Setting up detection conditions. If the input message does not meet the conditions, the input message is supplemented and combined. If the input message meets the conditions, the detection is performed directly. Step D: Use a sliding window mechanism to detect the transformed input message. If the input message contains illegal words, the user will be reminded and the message will be blocked from being sent.
2. According to claim 1, a harmful sentence identification method for the private chat function of the dating software is characterized in that: Also includes: Step E: Update the list of high-risk emojis based on customer feedback, and regularly update the frequency of use of each extended meaning of the emoji.
3. According to claim 1, a harmful sentence identification method for the private chat function of the dating software is characterized in that: The emoticon in step A has a unique corresponding code, and the code corresponding to the unique emoticon corresponds to multiple extended meanings.
4. According to claim 1, a harmful sentence identification method for the private chat function of the dating software is characterized in that: The step C comprises: Step C1, setting a detection condition of not less than four bytes; Step C2: if the input message does not meet the detection condition, the sent message is combined with the input message until the length of the input message that fails the detection meets the detection condition; Step C3: Detect the input message that meets the conditions.
5. According to claim 1, a harmful sentence identification method for the private chat function of the dating software is characterized in that: The step D comprises: Step D1, setting the sliding window to 6 coding units and setting the step length to one coding unit; Step D2, when the input message exceeds the capacity of the sliding window, the sliding window starts to perform sliding detection on the input message, starting from the starting position of the input message, and each sliding captures a character string of the window size, and each sliding is a step size; Step D3, each time the sliding window slides, a substring is generated in the string. If there is an emoticon in the substring, the corresponding meaning group is found according to the emoticon encoding, and after replacing the emoticon content with the meaning, semantic detection is performed on the replaced message; Step D4: If the sentence is detected to contain harmful information, the message is blocked from being sent, and the offending words in the input message are pointed out to the customer.
6. A method for identifying harmful sentences for the private chat function of a dating software according to claim 5, characterized in that: Also includes: Step D5, repeat steps D1-D4, and if no harmful information is detected, the message is allowed to be sent.
7. According to claim 1, a harmful sentence identification method for the private chat function of the dating software is characterized in that: The sliding window detection in step D is performed simultaneously with the user inputting the message, and the customer is reminded in real time.
8. According to claim 5, a method for identifying harmful sentences for the private chat function of a dating software, characterized in that: The step D3 further comprises: When performing meaning group replacement, the initial meaning of the emoticon is preferentially selected to be replaced in the sentence, and then the extended meaning is used for replacement and semantic detection is performed on the replaced message.
9. A method for identifying harmful sentences for the private chat function of a dating software according to claim 8, characterized in that: The replacement order of the extended meaning is determined according to the usage frequency probability value of the mark, and the extended meaning with a high probability value is replaced first.
Citation Information
Patent Citations
Expression robot applied to instant messaging tool
CN102750555A
Method and device for detecting words
CN102902766A
Bullet screen emotion analysis method and device
CN110569354A
Semantic analysis method based on emoji
CN110765300A
Retrieval method and device for sensitive words in text, electronic equipment and storage medium
CN111737398A