With the widespread adoption of various interactive systems, the security of
interactive content has become increasingly prominent. Users may output inappropriate remarks, and systems may generate inappropriate content, affecting user experience and even triggering legal risks. Existing
content security mechanisms suffer from deficiencies such as one-way protection, simple keyword matching, lack of tiered response, absence of self-reflection mechanisms, and lack of mediation mechanisms. Therefore, a two-way interactive security
protection mechanism is needed, which can effectively detect and process inappropriate content input by users, control inappropriate content output by the
system, reduce false positives through intent classification, and provide
flexible mechanisms such as self-reflection, third-party intervention, and cooling-off periods to achieve a healthy and safe interactive environment. This application provides an
interactive processing method,
system, and computer-readable storage medium, aiming to achieve two-way protection for both the generator and the
receiver, reduce false positives through intent classification, provide tiered response for progressive
processing, empower the generator with the ability to reflect and apologize, and facilitate third-party mediation.