Safety auditing method and device for interaction content between user and model, medium and product
By deploying interception and parsing of data packets on the API gateway, building user portraits, obtaining a set of target audit models, and performing double security audits on user request and model response data, we solve the problem that existing systems have difficulty identifying non-compliant content, and achieve a full security audit of the content of user and model interactions, ensuring the security and compliance of the content.
Patent Information
- Application Number
- CN202511195882.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-26
AI Technical Summary
The existing user-model interaction content security audit system is difficult to effectively identify and block non-compliant content, especially in the face of complex circumvention methods, which affects the stability and security of model services.
By deploying interception and parsing of data packets on the API gateway, building user portraits, obtaining a set of target audit models, and performing double security audits on user requests and model response data, the eBPF program is used to intercept data packets, and accurate audits are performed based on IP addresses and user behavior characteristics.
It achieves a full security audit of the interaction content between users and models, improves the accuracy and reliability of detection, avoids interference with intelligent models, ensures the security and compliance of content, and effectively responds to complex avoidance challenges.
Smart Images

Figure CN120729637A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of API content auditing, and in particular to a method, device, medium, and product for security auditing of content interacting between users and models. Background Art
[0002] With the rapid development of artificial intelligence (AI), large language models (LLMs) and large multimodal models have been widely used in numerous scenarios, including dialogue generation, content creation, and intelligent customer service. These models can provide users with efficient and intelligent services, significantly improving productivity in daily life. However, security and compliance issues surrounding the content of user interactions with these models have become increasingly prominent, attracting widespread attention from all sectors of society. Against this backdrop, technologies for auditing and filtering the input and output of large models have emerged.
[0003] However, most audit systems use intrusive auditing methods, which can disrupt the normal operation of model services and affect service performance and stability. Furthermore, with the continuous improvement of model capabilities and the continuous evolution of attack methods, content security audits face more complex and subtle circumvention challenges. Attackers attempt to bypass conventional content detection and filtering through context avoidance, exploiting semantic variations, homophonic words, internet buzzwords, and even combining multimodal input. These new circumvention methods significantly increase the difficulty of detection, making it difficult for existing content security systems to effectively identify and block non-compliant content, and unable to fully guarantee the content security and compliance of model service platforms. Summary of the Invention
[0004] In view of this, the embodiments of the present disclosure provide a security audit method, device, medium and product for user-model interaction content, which can intercept request data and response data at the API gateway to achieve dual security audit of user requests and model response data, avoid intrusive audit interference, effectively deal with complex circumvention methods, and ensure the content security and compliance of the model service platform.
[0005] In a first aspect, embodiments of the present disclosure provide a method for security auditing of user-model interaction content, employing the following technical solutions: Deploy an API gateway, intercept and parse data packets passing through the API gateway, obtain the user's IP address and the request data entered by the user, and determine the target intelligent model requested by the user; Building a user profile based on the IP address and the request data; Based on the IP address and the user profile, obtaining a target audit model set; Performing a security audit on the request data using the target audit model set; When the security audit of the request data passes, the request data is input into the target intelligent model, and the response data output by the target intelligent model is intercepted at the API gateway; A security audit is performed on the reply data, and when the reply data passes the security audit, the reply data is fed back to the user.
[0006] Optionally, intercepting and parsing data packets passing through the API gateway, obtaining the user's IP address and the user's request data, and determining the target intelligent model requested by the user includes: Mounting eBPF to a preset location and using eBPF to intercept all data packets passing through the API gateway; wherein the preset location includes but is not limited to the network interface layer, the entry point of the network protocol stack, and the network connection point of the API gateway; Parse the IP header of the data packet according to the IP protocol specification to obtain the user's IP address; Determine the application layer protocol used by the data packet by parsing the transport layer header of the data packet; Extracting a complete request body from the data packet based on the application layer protocol; Extracting the request data input by the user from the complete request body based on the application layer protocol; A model identifier is extracted from the header of the complete request body, and a target intelligent model is determined based on the model identifier.
[0007] Optionally, constructing a user profile based on the IP address and the request data includes: Obtain the user's tenant information by matching the IP address with a preset user information database; Obtain user's historical access behavior data; Preprocessing the request data, integrating the tenant information, the historical access behavior data, and the preprocessed request data into a user data set; Key features are extracted from the user dataset, and a user profile is constructed based on the key features.
[0008] Optionally, acquiring a target audit model set based on the IP address and the user profile includes: Use the IP address library to resolve the IP address and obtain the user's geographic location; determining a first audit model from a set of regional audit models based on the geographic location; Obtaining the user's occupation based on the IP address and the user profile; Determining a second audit model from a set of occupational audit models based on the user's occupation; Extracting user age from the user portrait; determining a third audit model from a set of age audit models based on the user's age; The first audit model, the second audit model and the third audit model are all target audit models, and all target audit models constitute a target audit model set.
[0009] Optionally, obtaining the user's occupation based on the IP address and the user portrait includes: Matching the IP address with all preset fields in a field database; wherein the field database stores mapping relationships between preset fields and occupation types; If the IP address contains any preset field, the occupation type corresponding to the preset field is used as the user's occupation; If the IP address does not contain any preset fields, the user occupation is extracted from the user portrait.
[0010] Optionally, performing a security audit on the request data using the target audit model set includes: Input the request data into each target audit model respectively, and obtain the sensitive category and corresponding confidence level output by each target audit model; Get the risk threshold for each sensitive category; Compare the confidence level of each sensitive category with the corresponding risk threshold; If the confidence level of any sensitive category is greater than or equal to the corresponding risk threshold, the security audit of the requested data is determined to have failed; If the confidence levels of all sensitive categories are less than the corresponding risk thresholds, it is determined that the security audit of the requested data has passed.
[0011] Optionally, obtaining the risk threshold for each sensitive category includes: Extracting user behavior feature labels and security feature labels from the user portrait; Obtaining a weight coefficient based on the behavior feature label and the security feature label; The weight coefficient is multiplied by the preset basic threshold of each sensitive category to obtain the risk threshold of each sensitive category.
[0012] Optionally, the security audit method for user-model interaction content further includes: When the security audit of the reply data fails, the request data and the reply data are combined into a set of related data and stored in a backtracking database; Each target audit model in the target audit model set is fine-tuned in a supervised manner based on all associated data in the backtracking database.
[0013] In a second aspect, the embodiments of the present disclosure further provide a security audit system for user-model interaction content, which employs the following technical solutions: The request data acquisition module is used to deploy the API gateway, intercept and parse the data packets passing through the API gateway, obtain the user's IP address and the request data entered by the user, and determine the target intelligent model requested by the user; A user portrait building module, configured to build a user portrait based on the IP address and the request data; An audit model acquisition module, configured to acquire a target audit model set based on the IP address and the user profile; A request data audit module, configured to perform a security audit on the request data using the target audit model set; A reply data interception module is used to input the request data into the target intelligent model when the security audit of the request data passes, and intercept the reply data output by the target intelligent model at the API gateway; The reply data audit module is used to perform a security audit on the reply data, and when the reply data security audit passes, the reply data is fed back to the user.
[0014] In a third aspect, the embodiments of the present disclosure further provide a computer device that adopts the following technical solution: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any of the above-mentioned security audit methods for user-model interaction content.
[0015] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the above-mentioned methods for security auditing of user-model interaction content.
[0016] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.
[0017] The security audit method for user-model interaction content provided by the embodiments of the present disclosure can intercept and parse the source of user-intelligent model interaction by deploying an API gateway at the front end of the large model inference service. User portraits can comprehensively reflect the characteristics and behavior patterns of users. By combining IP addresses and request data to construct user portraits, we can understand the user's usage habits, preferences, historical request content and other information. Different users may have different security risks, different IP addresses may come from different network environments, and user portraits also reflect the diverse characteristics of users. According to the IP address and user portrait, the target audit model set is obtained, and the target audit model that best suits the user can be selected to improve the accuracy and pertinence of the audit. The target audit model set contains multiple target audit models, each of which may audit the request data from different angles, greatly improving the accuracy and reliability of the detection. The request data is input into the target intelligent model only after it has passed the security audit, avoiding interference and potential attacks on the target intelligent model by non-compliant data. By intercepting the response data output by the target intelligent model at the API gateway, we achieve a comprehensive audit of user-model interactions. Even if the request data passes the audit, the target intelligent model may generate non-compliant responses during processing for various reasons. Auditing the response data further filters out potential security risks and non-compliant information, ensuring that the content ultimately fed back to users is secure and compliant. Furthermore, auditing the response data enhances the ability to identify hidden risks employed by attackers, such as context avoidance, exploitation of semantic variations, homophonic words, and internet buzzwords. This effectively addresses new evasion challenges and promotes the healthy development of large-scale model inference service platforms.
[0018] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the following specifically cites preferred embodiments and describes them in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A flowchart of a method for security auditing of user-model interaction content provided by an embodiment of the present disclosure; Figure 2A flowchart of a data packet interception and parsing method provided in an embodiment of the present disclosure; Figure 3 A flowchart of a method for constructing a user profile according to an embodiment of the present disclosure; Figure 4 A flowchart of a method for determining a target audit model provided by an embodiment of the present disclosure; Figure 5 A flowchart of a method for obtaining a user's occupation provided in an embodiment of the present disclosure; Figure 6 A flowchart of a method for requesting data security auditing provided in an embodiment of the present disclosure; Figure 7 A functional block diagram of a security audit system for user-model interaction content provided by an embodiment of the present disclosure; Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0022] It should be clear that the following embodiments of the present disclosure are described through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0023] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.
[0024] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0025] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.
[0026] Reference Figure 1 The present disclosure provides a security audit method for user-model interaction content, comprising the following steps: S1: Deploy an API gateway, intercept and parse data packets passing through the API gateway, obtain the user's IP address and the request data entered by the user, and determine the target intelligent model requested by the user; S2: Build user profiles based on IP addresses and request data; S3: Obtain the target audit model set based on IP address and user profile; S4: Use the target audit model set to perform security audit on the request data; S5: When the security audit of the request data passes, the request data is input into the target intelligent model, and the response data output by the target intelligent model is intercepted at the API gateway; S6: Perform a security audit on the reply data. When the reply data passes the security audit, the reply data will be fed back to the user.
[0027] This disclosure provides a security audit method for user-model interaction content. An API Gateway is a middleware component that acts as a "gatekeeper" for a system or service. By deploying the API Gateway at the front end of the large-scale model inference service, it can intercept and parse user interactions with intelligent models at the source, ensuring that all user request packets passing through the gateway are processed without omission. This fully covers all scenarios of user-model interaction and provides a complete data foundation for subsequent security audits. User profiles comprehensively reflect user characteristics and behavior patterns. By combining IP addresses and request data to construct user profiles, we can understand user usage habits, preferences, historical request content, and other information.
[0028] Different users may have different security risks, and different IP addresses may originate from different network environments. User profiles also reflect the diverse characteristics of users. By obtaining a target audit model set based on IP address and user profile, we can select the target audit model that best suits the user, ensuring that the target audit model matches the user's actual situation and improving the accuracy and pertinence of the audit. As user behavior changes and new security threats emerge, the target audit model set can be dynamically adjusted based on changes in user profiles and IP addresses, ensuring that audit policies remain effective.
[0029] The target audit model set includes multiple target audit models, each of which may audit request data from a different perspective. This multi-angle audit approach can more comprehensively detect security risks and non-compliant content in request data, greatly improving the accuracy and reliability of detection. Request data is only input into the target intelligent model after it has passed the security audit. This prevents interference and potential attacks from non-compliant data on the target intelligent model, ensures the normal operation and stability of the target intelligent model, and reduces the possibility of the target intelligent model malfunctioning, outputting erroneous results, or displaying sensitive information due to malicious data input.
[0030] By intercepting the response data output by the target intelligent model at the API gateway, a full audit of the interaction between the user and the model is implemented. This approach not only focuses on the security of user input but also audits the output of the target intelligent model, ensuring that the content of the entire interaction process meets security and compliance requirements. Even if the request data passes the audit, the target intelligent model may generate non-compliant response content during processing for various reasons. Auditing the response data can further filter out potential security risks and non-compliant information, ensuring that the content ultimately fed back to the user is safe and compliant. Furthermore, auditing the response data enhances the ability to identify hidden risks, such as context avoidance, semantic variants, homophonic words, and internet buzzwords, employed by attackers. This effectively addresses new evasion challenges and promotes the healthy development of large-model inference service platforms.
[0031] In S1, during the interaction between the user and the target intelligent model, the request data input by the user to the target intelligent model and the response data output by the target intelligent model together constitute the interaction content. To ensure the compliance of the interaction content, an API gateway is deployed at the front end of the large model inference service. The API gateway intercepts the request data input by the user so that it can be subsequently audited to ensure the security and legitimacy of the data. Figure 2The flowchart of the packet interception and parsing method shown here, "intercepting and parsing packets passing through the API gateway, obtaining the user's IP address and request data, and determining the target intelligent model for the user's request," includes the following steps: S11: Mount eBPF to a preset location and use eBPF to intercept all data packets passing through the API gateway; the preset location includes but is not limited to the network interface layer, the entry point of the network protocol stack, and the network connection point of the API gateway; S12: Parse the IP header of the data packet according to the IP protocol specification to obtain the user's IP address; S13: Determine the application layer protocol used by the data packet by parsing the transport layer header of the data packet; S14: extracting the complete request body from the data packet based on the application layer protocol; S15: Extracting the request data input by the user from the complete request body based on the application layer protocol; S16: Extract the model identifier from the header of the complete request body, and determine the target intelligent model based on the model identifier.
[0032] In S11, eBPF (extended Berkeley Packet Filter) is a mechanism for running user-defined programs in the Linux kernel. It offers high performance, zero intrusion, and secure isolation. It can be used to accurately capture request data in kernel mode without modifying user-mode business logic. Pre-written eBPF programs are loaded into the kernel using bpftool or other tools. The loading process verifies the eBPF program to ensure its security and correctness. The eBPF program is then mounted on key network processing paths within the operating system kernel, capturing network request data passing through the API gateway in real time without requiring any modifications to the existing business logic or model services. This is non-invasive and protocol-agnostic. These key network processing paths include, but are not limited to, the network interface layer, entry points into the network protocol stack, and network connection points of the API gateway. At the network interface layer, the tc (Traffic Control) tool can be used to mount eBPF programs onto the ingress or egress queues of a network interface. Here, the socket_filter mechanism can work in conjunction with eBPF programs. Socket_filter allows for filtering rules to be set at the socket level. When packets enter or leave the network interface, they are initially filtered by the socket_filter, after which the eBPF program is triggered to further capture and process the packets. At the entry points of the network protocol stack, eBPF programs can be mounted onto kernel functions using mechanisms such as kprobes or tracepoints. tcp_sendmsg is a key function in the Linux kernel network protocol stack for handling TCP data transmission. When the kernel processes a network packet and calls tcp_sendmsg, the eBPF program mounted there intercepts the packet and performs appropriate operations, such as parsing the IP header and transport layer header. Socket_filter can also be applied at this location to filter packets entering the protocol stack, ensuring that only packets that meet the filtering criteria are further processed. For the API gateway's network connection points, depending on the specific implementation of the API gateway, it may be necessary to use specific hooks or extension mechanisms to mount the eBPF program to the corresponding location. During this process, socket_filter can assist the eBPF program in filtering and controlling the data packets entering and leaving the API gateway, improving the API gateway's data processing efficiency and security.
[0033] In S12, after the eBPF program captures a packet, it needs to locate the starting position of the IP header. In an Ethernet frame, the IP header typically follows the Ethernet header, which is typically 14 bytes long. Therefore, the IP header can be found using an offset. According to the IP protocol specification, the IP header contains multiple fields, such as the version number, header length, source IP address, and destination IP address. The values of these fields can be obtained by reading the corresponding byte positions. For example, the source IP address in the IPv4 header is located in bytes 12-15. The eBPF program can read these four bytes and convert them into an IP address format. Considering both IPv4 and IPv6, the IPv6 header format differs from IPv4 and is longer. Therefore, the eBPF program needs to determine whether the packet is IPv4 or IPv6 based on the version number field in the IP header and apply the appropriate parsing method. The parsed source IP address is the user's IP address.
[0034] In S13, after parsing the IP header, the transport layer protocol (such as TCP or UDP) is determined based on the protocol field of the IP header, and then the starting position of the transport layer header is calculated based on the header length field of the IP header. For the TCP header, it contains fields such as the source port, destination port, and sequence number; for the UDP header, it contains fields such as the source port, destination port, and length. By reading the values of these fields, the source port and destination port information can be obtained. Different application layer protocols usually use fixed port numbers. For example, the HTTP protocol uses ports 80 or 443 by default, while some specific ports may be used for other application layer protocols such as gRPC services. The eBPF program can determine the application layer protocol used by the data packet based on the parsed port number.
[0035] In S14, once the application layer protocol is determined, the eBPF program extracts the request body from the data packet according to the corresponding protocol specifications. For HTTP, the starting position and length of the request body are found according to the HTTP request format, and the request body content is extracted. For gRPC, the gRPC request body is also extracted from the data packet according to its protocol rules.
[0036] The extracted request body data is stored in the RingBuffer. As an efficient circular buffer, the RingBuffer enables rapid data transmission from kernel mode to user mode, avoiding frequent data copying and significantly improving transmission efficiency. In user mode, the program reads the extracted request body data from the RingBuffer. Because network packets may be split into multiple fragments for transmission, session reassembly is required in user mode. For the HTTP protocol, the scattered packet fragments are reassembled into a complete HTTP request based on the HTTP request header information (such as the Content-Length field) and the request body data. For the gRPC protocol, session reassembly is performed according to its protocol rules to restore the complete gRPC request and ultimately obtain the complete request body.
[0037] In S15, in the HTTP protocol, request data may be distributed across different parts of the request body. When users use the model, POST requests are typically used, and the request data is typically in the request body. Depending on the format of the request body (e.g., form data, JSON data, etc.), the corresponding parsing method is used. For example, if the request body is in JSON format, a JSON parsing library (such as the json module in Python) can be used to parse it into a Python object, thereby extracting the user's request data. For gRPC requests, the corresponding serialization and deserialization methods are used to extract the request data based on the message structure defined by the protocol. gRPC typically uses ProtocolBuffers for data serialization, so the ProtocolBuffers library is required to parse the request body and obtain the user's request data.
[0038] In S16, the header information of the complete request body is searched for fields related to the model identifier. Different systems may use different field names to identify intelligent models. Common examples include custom request header fields such as X-Target-Model and Model-Name. The model identifier is extracted by reading the value of the corresponding header field. Based on the extracted model identifier, the corresponding intelligent model is searched in the model list of the large model inference service. This model is the target intelligent model requested by the user.
[0039] In S2, refer to Figure 3 The flowchart of the user profile construction method shown in the figure "Building a user profile based on IP address and request data" includes the following steps: S21: Obtain the user's tenant information by matching the IP address with the preset user information database; S22: Obtain the user's historical access behavior data; S23: Preprocess the request data and integrate tenant information, historical access behavior data, and preprocessed request data into a user data set; S24: Extract multiple key features from the user data set according to preset types, and build a user profile based on the key features.
[0040] In S21, a user information database is pre-set in the form of key-value pairs. This database stores a mapping between IP addresses and tenant information, with the IP address as the key and the corresponding tenant information (such as tenant ID, tenant name, tenant industry, tenant size, etc.) as the value. After obtaining the user's IP address, it is used as a query condition to search the pre-set user information database and extract the corresponding tenant information.
[0041] In S22, you can choose to use the tracking code provided by third-party statistical tools such as Google Analytics and Baidu Statistics, or you can develop your own tracking code based on business needs. The tracking code will be automatically executed when the user visits the page, and collect the user's access data, such as access time, page URL, length of stay, page jump path, etc. This data will be sent to the designated server for storage and analysis. On the server side, by writing code to record relevant access information in the process of processing user requests, you can use the log framework to record the user's IP address, request time, request method, request URL and other information in the log file, and extract the user's access information from the log file. Integrate the access data collected by the tracking code and the access information extracted from the log file into the user's historical access behavior data.
[0042] In S23, preprocessing includes data cleaning, formatting and standardization, where data cleaning refers to removing invalid data, duplicate data and noise data in the request data. For example, if the request data contains some special characters or garbled characters, they can be filtered and replaced; formatting refers to converting the request data into a unified format for subsequent analysis and processing, for example, converting date and time data into a standard format, encoding text data, etc.; standardization refers to standardizing the request data so that it has the same scale and range. For example, the normalization method can be used to scale numerical data to between 0 and 1.
[0043] Use database table join operations or data processing tools (such as the Pandas library) to merge tenant information, historical access behavior data, and preprocessed request data into a data structure to form a user dataset for subsequent feature extraction and analysis.
[0044] In S24, the types of key features include basic attribute features, behavioral features, and security features, among which basic attribute features include but are not limited to the user's age and occupation. Users of different age groups may have different risk preferences and common security issues when using large models. For example, teenagers may be more susceptible to negative information, and stricter content review may be required during audits; while the elderly may be more susceptible to fraudulent information, and during audits, attention should be paid to whether there are fraudulent responses. Different occupations involve different specific types of information and data. For example, financial practitioners may enter sensitive financial information, and during audits, they need to focus on data security and compliance; medical personnel may enter patients' private information, and they need to ensure that the information is not leaked.
[0045] Behavioral characteristics include, but are not limited to, the user's historical request frequency, request time patterns, and request content types. Frequent requests may indicate abnormal behavior, such as malicious attacks or batch data acquisition, so a high historical request frequency requires focused auditing. Whether a user's request time is regular can also reflect their behavior patterns. For example, a user who usually uses a large model during daytime working hours but suddenly makes frequent requests late at night may be experiencing anomalies and require a more rigorous audit. The types of content frequently requested by users, such as text, images, videos, etc., as well as specific subject areas such as technology, entertainment, health, etc., can all provide references for security audits. Different types of content may present different security risks. For example, requests involving politically sensitive topics or personal privacy information require special attention.
[0046] Security features include, but are not limited to, the number of violations a user has committed in the past and information about the device they are using. If a user has committed multiple violations in the past, such as entering malicious code or spreading false information, these records will become important references for security audits. For users with a history of violations, stricter audit policies will be adopted. Information such as the device type and operating system version used by the user is also relevant to security audits. For example, devices using outdated operating systems may have more security vulnerabilities, requiring more stringent checks on their requested data.
[0047] To ensure that key features accurately reflect user characteristics and effectively serve subsequent analysis and decision-making, users are assigned multiple targeted tags by matching key features with pre-set tagging rules. For example, if a user is 25 years old, which falls within the pre-set age range of 19-35, they are labeled "youth." If a user initiates more than 100 requests in a day, they are labeled "high-frequency request" according to the rules. By integrating these key features and the user tags generated by matching them, a user profile is constructed that comprehensively and meticulously depicts user characteristics and behavior patterns.
[0048] In S3, refer to Figure 4 The flowchart of the target audit model determination method shown in the figure, "Obtaining a target audit model set based on IP address and user profile," includes the following steps: S31: Use the IP address library to resolve the IP address and obtain the user's geographic location; S32: Determine a first audit model from a set of regional audit models based on the geographical location; S33: Obtain user occupation based on IP address and user profile; S34: Based on the user's occupation, determine a second audit model from the occupation audit model set; S35: Extract user age from user portrait; S36: Determine a third audit model from the age audit model set based on the user's age; S37: The first audit model, the second audit model and the third audit model are all target audit models, and all target audit models constitute a target audit model set.
[0049] In S31, an IP address library such as GeoIP or IPIP.net is selected, and according to the API or tool provided by the selected IP address library, the user's IP address is used as input, and the corresponding parsing function or interface is called to obtain the user's geographic location information, including country, province, city, region, etc.
[0050] Reference Figure 5 The flowchart of the method for obtaining user occupations is shown in the figure. "Acquiring user occupations based on IP address and user profile" includes the following steps: S331: Match the IP address with all preset fields in the field database; if the IP address contains any preset field, execute S332; if the IP address does not contain any preset field, execute S333; S332: The occupation type corresponding to the preset field is used as the user's occupation; S333: Extract user occupation from user portrait.
[0051] The field database stores mappings between preset fields and occupation types, with the preset fields serving as keys and the corresponding occupation types serving as values. IP addresses are readily accessible, critical information during online interactions. By matching IP address-related information with preset fields, a user's occupation can be quickly and preliminarily determined. This approach avoids complex analysis of user profiles and significantly improves data processing efficiency. In terms of data timeliness, IP addresses reflect a user's current online environment in real time and can be directly linked to the user's company or organization. This makes occupation information determined based on IP addresses more relevant to the user's current situation and more accurate. In contrast, occupation information in user profiles is often derived from past user input. As users' occupations may change over time, this introduces the risk of information lag. Based on these characteristics, when obtaining user occupation information, IP address-related information is prioritized, leveraging its timeliness and relevance. Only when the user's occupation is difficult to determine based on the IP address should relevant information be extracted from the user profile. This creates a more efficient and accurate user occupation determination process. To enhance the timeliness of user profiles, user profiles must be promptly updated when user information changes.
[0052] In content security auditing and filtering scenarios, due to varying laws and regulations across countries and regions, as well as the unique compliance requirements of different tenants (enterprises and institutions), content security auditing and filtering standards vary significantly. Different industries or enterprises also develop their own customized content compliance standards. Different age groups may have different content access restrictions and different auditing rigor. Existing content auditing systems that fail to adapt to these multi-dimensional differences will struggle to meet compliance requirements across countries, regions, industries, and age groups, potentially leading to inaccurate audit results or non-compliance with the requirements of specific regions and tenants. This solution reflects the user's country and region based on their geographic location and selects a corresponding first audit model from a set of regional audit models. This allows the API gateway to match the user's country or region with the applicable audit rules and standards. This solution also selects a suitable second audit model based on the user's occupation and a corresponding third audit model based on age. These audit models are then used to perform security audits on request data, ensuring audit rules are more tailored to the user's specific circumstances. This allows for more accurate identification and filtering of content that does not comply with local laws and regulations, industry standards, and the enterprise's personalized requirements, avoiding misjudgments or omissions that can occur with single-standard audits.
[0053] In S4, refer to Figure 6 The flowchart of the request data security audit method shown in the figure, "Security audit of request data using a target audit model set," includes the following steps: S41: Input the request data into each target audit model respectively, and obtain the sensitive category and confidence level of the sensitive category output by each target audit model; S42: Obtain the risk threshold for each sensitive category; S43: Compare the confidence level of each sensitive category with the corresponding risk threshold. If the confidence level of any sensitive category is greater than or equal to the corresponding risk threshold, execute S44. If the confidence levels of all sensitive categories are less than the corresponding risk threshold, execute S45. S44: determining that the security audit of the requested data has failed; S45: Determine whether the security audit of the requested data has passed.
[0054] In the security audit process described above, each target audit model conducts a detailed analysis of the input request data. The target audit model accurately identifies the sensitive information within the request data, specifies the specific type of this sensitive information, and then outputs the corresponding sensitive category and the confidence level for that sensitive category. The confidence level reflects the model's confidence in the data's belonging to a specific sensitive category. The confidence levels for these sensitive categories are then compared with pre-set risk thresholds, which are based on the potential security risks posed by different sensitive categories. This comparison enables a scientific and accurate determination of the security of the request data. Only when the confidence levels for all sensitive categories are below the corresponding risk thresholds is the security audit for the request data considered passed. This approach provides a comprehensive, detailed, and accurate security review of the request data, effectively identifying potential sensitive information and its level of risk.
[0055] Furthermore, the user's behavioral feature labels and security feature labels are extracted from the user profile. Based on these behavioral feature labels and security feature labels, a weight coefficient is obtained. This weight coefficient is multiplied by the preset basic threshold of each sensitive category to obtain the risk threshold of each sensitive category. The calculation formula of the weight coefficient is: Where, represents the weight coefficient; Indicates the The weight score of the behavioral feature label; Indicates the total number of behavioral feature labels; Indicates the The weight scores of the security feature labels; Indicates the total number of security feature tags.
[0056] Among them, the behavioral feature label includes at least one of the historical request frequency label, the request time pattern label and the request content type label; the historical request frequency label includes one of the high-frequency request label, the medium-frequency request label and the low-frequency request label; the request time pattern label includes one of the abnormal request time label and the normal request time label; the request content type label is an identifier that summarizes and classifies the key features and attributes of the request content. The content it covers will vary depending on different application scenarios and classification systems. For example, it includes text, image, audio and video categories, and is further subdivided into more specific labels under different categories. For example, the video category is divided into movie labels, TV series labels, short video labels and teaching video labels.
[0057] Different tags are assigned different weight scores. Taking the request frequency tag as an example, the weight score of the high-frequency request tag is set to a positive number, and the weight score of the low-frequency request tag is set to a negative number. By associating the obtained user tags with their corresponding weight scores and summing these weight scores, the weight coefficient applicable to the current user can be obtained. By using this weight coefficient to adjust the basic threshold, a more reasonable risk threshold can be obtained. This method of dynamically adjusting the risk threshold based on user tags, on the one hand, ensures the comprehensiveness and accuracy of data security audits and effectively protects data security; on the other hand, it avoids the problem of excessive review caused by the use of unified and fixed audit standards, significantly improves audit efficiency, and enables the system to achieve efficient and stable operation on the basis of meeting security compliance requirements.
[0058] In S5 and S6, simply conducting a security audit on the user's request data is insufficient to fully assess potential security risks throughout the entire interaction process. This is because the request data may only be surface information. During data processing by the target intelligent model, the output response data may contain security risks, such as sensitive information leakage, misleading information, or malicious induction, due to the model's inherent characteristics, limitations of the training data, or potential vulnerabilities. Therefore, even after the security audit of the request data passes, a security audit of the response data output by the target intelligent model is still required to ensure that the information ultimately presented to the user is secure, accurate, and compliant. This prevents inappropriate content in the response data from negatively impacting users, organizations, or society, further strengthening the security and reliability of the entire interaction process and ensuring the system operates securely and stably. When conducting a security audit on the response data, the same principles as the request data security audit can be employed, using multiple different types of security audit models. Alternatively, a unified security audit model can be used to obtain the corresponding sensitive categories and confidence levels, thereby determining the security audit results for the response data. Only when the response data security audit passes will the response data be intercepted at the API gateway and fed back to the user.
[0059] Furthermore, if the security audit of the response data fails, indicating an error in the previous security audit of the request data, the request data and the response data are combined into a set of related data and stored in a backtracking database. Based on all the related data in the backtracking database, each target audit model in the target audit model set is fine-tuned in a supervised manner. The backtracking database acts like a "problem data warehouse," specifically storing audited interaction data with security issues. Fine-tuning the target audit model in a supervised mode enables better identification of similar security issues. Over time, as problem data accumulates, the target audit model becomes increasingly intelligent, more effectively safeguarding the security of user input data.
[0060] A decision tree or rule chain is constructed. When the security audit of the request data fails, or when the security audit of the request data passes but the security audit of the reply data fails, the confidence level obtained during the security audit is matched with the decision tree or rule chain to determine the appropriate handling strategy, such as disconnection, rate limiting response, logging, and alarm reporting. The eBPF Map is a data structure provided by eBPF, similar to key-value storage, that can be used to share data between kernel state and user state. Therefore, policy decision results can be written back to the eBPF Map channel through user state and passed to kernel state. The deployed eBPF program reads the control signal and performs actual operations such as packet interception, RST response, and rate limiting. Policy control can be executed without leaving the kernel path, with latency as low as microseconds, achieving millisecond-level policy feedback. This solution builds a real-time response closed loop, avoiding the security failure problem of "slow recognition and slow feedback."
[0061] This solution's content identification and policy control functions independently from the model service, offering exceptional flexibility and real-time performance. Its value proposition is significant, as it is plug-and-play compatible with LLM or API services deployed in any language, framework, or platform. It supports local deployment, edge node operation, and cross-platform migration, resulting in high reusability. Furthermore, it supports hot updates of policy rules and models, combines multiple types of audit output, and integrates with cloud platforms like Kubernetes for multi-tenant policy distribution. It is suitable for AI service platforms or API gateways with high security requirements and complex traffic.
[0062] Reference Figure 7 The present disclosure provides a security audit system for user-model interaction content, including: The request data acquisition module 101 is used to deploy the API gateway, intercept and parse the data packets passing through the API gateway, obtain the user's IP address and the request data entered by the user, and determine the target intelligent model requested by the user; User profile building module 102, for building a user profile based on the IP address and request data; The audit model acquisition module 103 is used to obtain a target audit model set based on the IP address and user profile; The request data audit module 104 is used to perform security audit on the request data using the target audit model set; The reply data interception module 105 is used to input the request data into the target intelligent model when the security audit of the request data passes, and intercept the reply data output by the target intelligent model at the API gateway; The reply data audit module 106 is used to perform a security audit on the reply data. When the reply data security audit passes, the reply data is fed back to the user.
[0063] The various variations and specific examples of the security audit method for user-model interaction content provided above are also applicable to the security audit system for user-model interaction content provided in the present disclosure. Through the above detailed description of the security audit method for user-model interaction content, those skilled in the art can clearly know the implementation method of the security audit system for user-model interaction content. For the sake of brevity of the specification, it will not be described in detail here.
[0064] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc.
[0065] The processor can be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and can control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is configured to execute the computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the security audit method for user-model interaction content described in various embodiments of the present disclosure.
[0066] Those skilled in the art should understand that in order to solve the technical problem of how to obtain a good user experience, this embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the scope of protection of this disclosure.
[0067] like Figure 8The present invention provides a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 8 The computer device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0068] like Figure 8 As shown, a computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) or programs loaded from a storage device into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0069] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes and hard disks; and communication devices. The communication device can allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 8 A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0070] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the security audit method for user-model interaction content of the embodiment of the present disclosure are executed.
[0071] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.
[0072] According to an embodiment of the present disclosure, a computer-readable storage medium stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the security audit method for user-model interaction content described in each embodiment of the present disclosure are performed.
[0073] The above-mentioned computer-readable storage media include, but are not limited to, optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or mobile hard disks), media with built-in rewritable non-volatile memory (e.g., memory cards), and media with built-in ROM (e.g., ROM cartridges).
[0074] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.
[0075] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0076] In the present disclosure, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0077] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0078] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0079] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same function or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.
[0080] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0081] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A security audit method for user-model interaction content, characterized in that: include: Deploy an API gateway, intercept and parse data packets passing through the API gateway, obtain the user's IP address and the request data entered by the user, and determine the target intelligent model requested by the user; Building a user profile based on the IP address and the request data; Based on the IP address and the user profile, obtaining a target audit model set; Performing a security audit on the request data using the target audit model set; When the security audit of the request data passes, the request data is input into the target intelligent model, and the response data output by the target intelligent model is intercepted at the API gateway; A security audit is performed on the reply data, and when the reply data passes the security audit, the reply data is fed back to the user.
2. The security audit method for user-model interaction content according to claim 1 is characterized in that: The interception and analysis of the data packets passing through the API gateway, obtaining the user's IP address and the user's request data, and determining the target intelligent model requested by the user include: Mounting eBPF to a preset location and using eBPF to intercept all data packets passing through the API gateway; wherein the preset location includes but is not limited to the network interface layer, the entry point of the network protocol stack, and the network connection point of the API gateway; Parse the IP header of the data packet according to the IP protocol specification to obtain the user's IP address; Determine the application layer protocol used by the data packet by parsing the transport layer header of the data packet; Extracting a complete request body from the data packet based on the application layer protocol; Extracting the request data input by the user from the complete request body based on the application layer protocol; A model identifier is extracted from the header of the complete request body, and a target intelligent model is determined based on the model identifier.
3. The security audit method for user-model interaction content according to claim 1 is characterized in that: The constructing of a user profile based on the IP address and the request data includes: Obtain the user's tenant information by matching the IP address with a preset user information database; Obtain user's historical access behavior data; Preprocessing the request data, integrating the tenant information, the historical access behavior data, and the preprocessed request data into a user data set; Key features are extracted from the user dataset, and a user profile is constructed based on the key features.
4. The security audit method for user-model interaction content according to claim 1 is characterized in that: The acquiring of a target audit model set based on the IP address and the user profile includes: Use the IP address library to resolve the IP address and obtain the user's geographic location; determining a first audit model from a set of regional audit models based on the geographic location; Obtaining the user's occupation based on the IP address and the user profile; Determining a second audit model from a set of occupational audit models based on the user's occupation; Extracting user age from the user portrait; determining a third audit model from a set of age audit models based on the user's age; The first audit model, the second audit model and the third audit model are all target audit models, and all target audit models constitute a target audit model set.
5. The security audit method for user-model interaction content according to claim 4 is characterized in that: The obtaining of the user's occupation based on the IP address and the user portrait includes: Matching the IP address with all preset fields in a field database; wherein the field database stores mapping relationships between preset fields and occupation types; If the IP address contains any preset field, the occupation type corresponding to the preset field is used as the user's occupation; If the IP address does not contain any preset fields, the user occupation is extracted from the user portrait.
6. The security audit method for user-model interaction content according to claim 4 is characterized in that: The performing a security audit on the request data using the target audit model set includes: Input the request data into each target audit model respectively, and obtain the sensitive category and corresponding confidence level output by each target audit model; Get the risk threshold for each sensitive category; Compare the confidence level of each sensitive category with the corresponding risk threshold; If the confidence level of any sensitive category is greater than or equal to the corresponding risk threshold, the security audit of the requested data is determined to have failed; If the confidence levels of all sensitive categories are less than the corresponding risk thresholds, it is determined that the security audit of the requested data has passed.
7. The security audit method for user-model interaction content according to claim 6 is characterized in that: Obtaining the risk threshold for each sensitive category includes: Extracting user behavior feature labels and security feature labels from the user portrait; Obtaining a weight coefficient based on the behavior feature label and the security feature label; The weight coefficient is multiplied by the preset basic threshold of each sensitive category to obtain the risk threshold of each sensitive category.
8. The security audit method for user-model interaction content according to claim 4 is characterized in that: When the security audit of the reply data fails, the request data and the reply data are combined into a set of related data and stored in a backtracking database; Each target audit model in the target audit model set is fine-tuned in a supervised manner based on all associated data in the backtracking database.
9. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the security audit method for user-model interaction content described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the security audit method for user-model interaction content described in any one of claims 1-8.
11. A computer program product comprising computer instructions, characterized in that When the computer instruction is executed by a processor, the steps of the method for security auditing of user-model interaction content described in any one of claims 1-8 are implemented.
Citation Information
Patent Citations
Resource transaction auditing method and device and server
CN111523849A
Internet-oriented security auditing method and internet-oriented security auditing system
CN111683107A
Campus network security risk terminal interception traceability system
CN116781380A
API gateway log collecting and reporting method, system and device and medium
CN120151188A
Method and system for streamlined auditing
US20200167387A1