Information acquisition and analysis method, device, equipment, medium and product
Through the gateway and dynamic traffic processing platform with multi-protocol identification function, monitoring and forwarding LLM call requests, combining user behavior data for full-link multi-dimensional analysis, solving the problem of LLM call data dispersed and single dimensions in the prior art, achieving more accurate analysis results and higher credibility.
Patent Information
- Application Number
- CN202510600459.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
It is difficult for the prior art to realize unified management and comprehensive analysis of large language model (LLM) call data, especially when data is scattered, single acquisition dimensions and ignoring user behavior data.
Through a gateway with multi-protocol recognition function and a dynamic traffic processing platform, language model call requests transmitted by multiple data sources are monitored and forwarded to the language model service cluster to generate question and answer data. At the same time, the user behavior data generated by the target data source is obtained based on the Q&A data, and a full-link multi-dimensional analysis is performed.
It realizes unified management and comprehensive analysis of LLM call data, takes into account user behavior data, and the generated call analysis results are more accurate, which improves the credibility of the application effect of LLM.
Smart Images

Figure CN120106232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing technology, and in particular to an information collection and analysis method, device, equipment, medium and product. Background Art
[0002] With the rapid advancement of artificial intelligence technology, the application scenarios of Large Language Model (LLM) in and outside the enterprise are becoming more and more extensive, covering multiple fields such as Integrated Development Environment (IDE), User Interface (UI) and Application Programming Interface (API) calls. In order to better evaluate and optimize the performance of LLM, understand user interaction patterns, and perform fine management and resource scheduling, comprehensive collection and analysis of LLM call information has become crucial.
[0003] However, since the call information of different entrances is scattered across various systems and lacks a unified collection mechanism, it is difficult to centrally manage data and cross-entry analysis becomes complicated and difficult. Secondly, existing collection technologies often only focus on requests and responses, but ignore key information such as user behavior data, such as the adoption of code completion or UI interaction behavior, which limits the in-depth analysis of the application effect of LLM.
[0004] In view of the above, how to solve the current problem that LLM call data is scattered and not conducive to management, the collection dimension is single and ignores user behavior data, and it is difficult to achieve comprehensive, unified and efficient collection and analysis is an urgent problem that needs to be solved by technical personnel in this field. Summary of the invention
[0005] The present invention provides an information collection and analysis method to at least solve the problems that the current LLM call data is scattered and not conducive to management, the collection dimension is single and ignores user behavior data, and it is difficult to achieve comprehensive, unified and efficient collection and analysis.
[0006] The present invention provides an information collection and analysis method, comprising:
[0007] Monitoring language model call requests transmitted by multiple data sources through a gateway with multi-protocol identification function and a dynamic traffic processing platform; wherein the data sources include at least a user interface, an application programming interface, and an integrated development environment plug-in;
[0008] When receiving a language model call request transmitted by the target data source, forwarding the language model call request to the language model service cluster to generate corresponding question and answer data through the language model service cluster;
[0009] Forward the question and answer data to the target data source, and obtain the user behavior data generated by the target data source based on the question and answer data;
[0010] Perform full-link multi-dimensional analysis based on question-and-answer data and corresponding user behavior data to generate call analysis results for the language model service cluster.
[0011] The present invention also provides an information collection and analysis device, comprising:
[0012] A monitoring module, used to monitor language model call requests transmitted by multiple data sources through a gateway with multi-protocol identification function and a dynamic traffic processing platform; wherein the data source includes at least a user interface, an application programming interface, and an integrated development environment plug-in;
[0013] A forwarding module, configured to forward the language model call request to the language model service cluster when receiving the language model call request transmitted by the target data source, so as to generate corresponding question and answer data through the language model service cluster;
[0014] An acquisition module is used to forward the question and answer data to a target data source and obtain user behavior data generated by the target data source based on the question and answer data;
[0015] The analysis module is used to perform full-link multi-dimensional analysis based on question and answer data and corresponding user behavior data to generate call analysis results for the language model service cluster.
[0016] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any one of the above-mentioned information collection and analysis methods when executing the computer program.
[0017] The present invention also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned information collection and analysis methods are implemented.
[0018] The present invention also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned information collection and analysis methods when executed by a processor.
[0019] The information collection and analysis method provided by the present invention monitors the language model call requests transmitted by multiple data sources through a gateway with multi-protocol identification function and a dynamic traffic processing platform, thereby realizing unified management of LLM call data and adapting to multiple protocols; when a language model call request is received, the language model call request is forwarded to a language model service cluster, corresponding question and answer data is generated, and user behavior data generated by a target data source based on the question and answer data is obtained, so as to facilitate full-link multi-dimensional analysis based on the question and answer data and the corresponding user behavior data. The analysis process not only focuses on the call request and response, but also takes into account the corresponding user behavior, so that the generated call analysis result is more accurate, and the credibility of the LLM application effect analysis is improved.
[0020] In addition, the present invention also provides an information collection and analysis device, equipment, medium and product, with the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0022] Figure 1 A flowchart of an information collection and analysis method provided by an embodiment of the present invention;
[0023] Figure 2 An architecture diagram of an information collection and analysis system provided by an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of an information collection and analysis device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] It should be noted that, in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0027] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0028] With the rapid development of artificial intelligence technology, the application scenarios of LLM are becoming more and more diverse, covering IDE, UI, API calls and other aspects. In IDE plug-ins, developers frequently use LLM for code completion or Q&A to improve development efficiency; in UI, users can communicate and Q&A with LLM through web pages or applications to obtain intelligent interactive experience; and through API calls, the system or third-party applications can directly use LLM services to achieve more extensive functional integration.
[0029] However, current technologies face many challenges in the field of LLM call information collection. Since call information from different entrances is scattered across various systems and there is a lack of a unified collection mechanism, data is difficult to centrally manage, and cross-entry analysis becomes complicated and difficult. At the same time, existing collection technologies often only focus on requests and responses, but ignore key information such as user behavior, such as the adoption of code completion or UI interaction behavior, which limits in-depth analysis of the application effect of LLM. These problems jointly restrict the comprehensive, unified, efficient and intelligent collection and analysis of LLM call information, making it difficult to meet the urgent needs of performance evaluation and optimization. Therefore, in order to solve the above problems, the present invention provides an information collection and analysis method. It can be understood that the method provided by the present invention is applied to a computing device such as a server or a host, which is communicatively connected to multiple entry data sources (such as UI, API, and IDE plug-in) and is communicatively connected to an LLM service cluster.
[0030] Figure 1 The following is a flow chart of an information collection and analysis method provided by an embodiment of the present invention. Figure 1 As shown, the method includes:
[0031] S10: Monitor language model call requests transmitted by multiple data sources through a gateway with multi-protocol identification capabilities and a dynamic traffic processing platform.
[0032] The data source includes at least a user interface, an application programming interface, and an integrated development environment plug-in.
[0033] Specifically, the present invention pre-builds a multi-protocol access gateway, in which the gateway and dynamic traffic processing platform are used as the basic software platform, and has a multi-protocol identification function. In the specific implementation, the gateway and the dynamic traffic processing platform can use high-performance Web server software, such as open source software OpenResty or Nginx as the basic software platform of the multi-protocol access gateway, and use its high-performance, high-concurrency, and scalable characteristics to identify the traffic transmitted by different data sources through corresponding protocols, including but not limited to Hypertext Transfer Protocol version (HTTP), WebSocket protocol, and Server-Sent Events (SEE) protocol traffic, and identify the language model call requests therein. In this embodiment, there is no restriction on the identification method of the language model call request.
[0034] It should be noted that the data source includes at least UI, API and IDE plug-in, and may also include other data sources, which are not limited in this embodiment. In addition, the language model call request is a request to call the LLM service cluster to perform a task, including but not limited to question-and-answer tasks and code completion tasks, etc., depending on the specific implementation. It should also be noted that the types of language model call requests sent by different data sources may be the same or different.
[0035] S11: When a language model call request transmitted by a target data source is received, the language model call request is forwarded to the language model service cluster to generate corresponding question and answer data through the language model service cluster.
[0036] When the language model call request transmitted by the target data source is confirmed, the language model call request is forwarded to the language model service cluster so that the language model service cluster can generate the corresponding question and answer data according to the language model call request. For example, when the task in the language model call request is a question, the language model service cluster will generate the answer corresponding to the question; when the task in the language model call request is a code completion task, the language model service cluster will generate the completion code corresponding to the code completion task.
[0037] S12: Forward the question and answer data to a target data source, and obtain user behavior data generated by the target data source based on the question and answer data.
[0038] After the language model service cluster generates the question and answer data corresponding to the language model call request, the question and answer data is forwarded to the target data source, and the user behavior data generated by the target data source based on the question and answer data is obtained. It can be understood that the user behavior data is the behavior generated by the user based on the question and answer data after receiving the question and answer data. For example, the language model call request of the IDE plug-in to the language model service cluster includes a code completion task, and the language model service cluster will generate the corresponding completion code according to the code completion task; when the IDE plug-in receives the completion code output by the language model service cluster, the user can choose to adopt the completion code or refuse to adopt the completion code, and the adoption / rejection is the user behavior corresponding to the completion code. In this embodiment, there is no restriction on the specific method of obtaining the user behavior data generated by the target data source based on the question and answer data.
[0039] S13: Perform full-link multi-dimensional analysis based on the question-and-answer data and the corresponding user behavior data to generate call analysis results for the language model service cluster.
[0040] Finally, a full-link multi-dimensional analysis is performed based on the question-and-answer data and the corresponding user behavior data. The analysis process not only focuses on the call request and response, but also takes into account the corresponding user behavior, so that the generated call analysis results are more accurate. In this embodiment, there is no restriction on the specific process of performing a full-link multi-dimensional analysis based on the question-and-answer data and the corresponding user behavior data, and a variety of analysis methods can be used, depending on the specific implementation situation.
[0041] In this embodiment, a gateway with multi-protocol identification function and a dynamic traffic processing platform are used to monitor language model call requests transmitted by multiple data sources, thereby realizing unified management of LLM call data and adapting to multiple protocols. When a language model call request is received, the language model call request is forwarded to the language model service cluster, corresponding question and answer data is generated, and user behavior data generated by the target data source based on the question and answer data is obtained, so as to facilitate full-link multi-dimensional analysis based on the question and answer data and the corresponding user behavior data. The analysis process not only focuses on the call request and response, but also takes into account the corresponding user behavior, so that the generated call analysis results are more accurate, thereby improving the credibility of the LLM application effect analysis.
[0042] Figure 2 The architecture diagram of the information collection and analysis system provided by the embodiment of the present invention is shown in FIG. Figure 2 As shown, in order to monitor language model call requests, in some embodiments, a gateway with multi-protocol identification function and a dynamic traffic processing platform are used to monitor language model call requests transmitted by multiple data sources, including:
[0043] S101: Monitoring the traffic data transmitted by each data source through the corresponding protocol based on the gateway and the dynamic traffic processing platform.
[0044] S102: Identify language model call requests that meet language model call features in each flow data through deep packet inspection technology.
[0045] In order to be able to identify language model call requests, this embodiment is based on the gateway and dynamic traffic processing platform to monitor the traffic data transmitted by each data source through the corresponding protocol, and use deep packet inspection (DPI) technology to identify language model call requests that meet the language model call characteristics in each traffic data.
[0046] It should be noted that the process of DPI technology identifying language model call requests in traffic data involves in-depth analysis of data packets. When a data packet passes through a network device, the DPI system extracts and examines the header and payload information of the data packet. It uses predefined rules, feature libraries, or machine learning algorithms to identify traffic patterns for specific applications or services. For example, by analyzing the URL, User-Agent string, or characteristics of a specific protocol in an HTTP request, DPI can determine whether the request comes from a web browser, email client, or other application. In addition, DPI can also identify specific patterns or metadata in encrypted traffic to infer its content or purpose. In this way, DPI technology is able to accurately identify and classify language model call requests in various network traffic.
[0047] Specifically, the gateway and dynamic traffic processing platform use DPI technology to intelligently identify requests in various protocols that meet the model call characteristics. For example, multi-dimensional intelligent identification of requests that meet the model call characteristics is performed through URL path characteristics, request header characteristics, request body characteristics, etc., to improve recognition accuracy and reduce misjudgments.
[0048] On the basis of the above embodiment, in order to accurately forward the language model call request to the LLM service cluster, in some embodiments, when a language model call request transmitted by the target data source is received, the language model call request is forwarded to the language model service cluster, including:
[0049] S103: Generate a request fingerprint of the language model call request by using a preset identifier generation algorithm.
[0050] The request fingerprint includes at least metadata information of the user's Internet Protocol address, client type, entry type, model version and request timestamp.
[0051] S104: Add the request fingerprint to the request header of the language model call request.
[0052] S105: Forward the language model call request to the language model service cluster according to the request fingerprint in the corresponding request header.
[0053] Specifically, a request fingerprint generation module is developed in advance in the multi-protocol access gateway based on Lua script (a lightweight, efficient and embeddable scripting language), and the module can generate a request fingerprint of a language model call request through a preset identifier generation algorithm at runtime. For example, a request fingerprint of a language model call request is generated using a universally unique identifier (Universally Unique Identifier Generation Algorithm, UUID). It should be noted that UUID is a 128-bit number, usually represented by 32 hexadecimal characters, with extremely high uniqueness, and it is almost guaranteed that there will be no duplication. When generating a request fingerprint, a UUID can be created when the request is initiated and sent as a parameter or header information of the request. For example, in an HTTP request, the UUID can be attached to the request as a custom header (such as X-Request-ID). This allows each request to have a unique identifier, which is convenient for tracking, debugging and logging within the system. In addition, the request fingerprint includes at least metadata information of the user's Internet Protocol address, client type, entry type, model version and request timestamp, and may also include other metadata information, which is not limited in this embodiment.
[0054] Subsequently, the request fingerprint is added to the request header of the language model call request, thereby generating a language model call request containing the request fingerprint. Finally, the language model call request is forwarded to the LLM service cluster according to the request fingerprint in the corresponding request header. As a key identifier for full-link data association, the request fingerprint can solve the problems of complex heterogeneous protocol adaptation and lack of full-link tracking, ensuring the complete transmission of data.
[0055] In some embodiments, in order to forward the language model call request to the language model service cluster, forwarding the language model call request to the language model service cluster according to the request fingerprint in the corresponding request header includes:
[0056] S106: Configure the proxy delivery instructions between the gateway and the dynamic traffic processing platform.
[0057] S107: Based on the proxy delivery instruction, the control gateway and the dynamic traffic processing platform forward the language model call request to the language model service cluster according to the request fingerprint, entry type and model version.
[0058] Specifically, the proxy_pass instructions of the gateway and the dynamic traffic processing platform are configured. Based on the proxy_pass instructions, the gateway and the dynamic traffic processing platform are controlled to flexibly route and forward the language model call requests to the backend LLM service cluster according to the request fingerprint, entry type and model version, thereby achieving load balancing.
[0059] In summary, the multi-protocol access gateway uses the gateway and dynamic traffic processing platform as the basic software platform. It is the unified traffic entrance of the system and is responsible for functions such as protocol adaptive identification, multi-protocol access, request interception and forwarding, and dynamic request fingerprint generation. It solves the problems of complex heterogeneous protocol adaptation and lack of full-link tracking, and ensures the integrity and real-time performance of streaming data. In addition, in order to ensure the normal operation of the gateway and dynamic traffic processing platform, the configured gateway and dynamic traffic processing platform can also be deployed to an independent server or containerized environment for performance testing and functional verification to ensure the high performance, high reliability, and high availability of the gateway.
[0060] It can be seen from the above embodiments that the data source includes at least UI, API and IDE plug-in. Whether each data source will trigger behavior after receiving the question and answer data of the LLM service cluster mainly depends on the design purpose and system permissions. Specifically, the UI and API generally only display and output the question and answer data, and do not generate user behavior, while the IDE plug-in may generate corresponding user behavior based on the question and answer data (such as code data), such as adopting or not adopting the code. Therefore, in some embodiments, the question and answer data is forwarded to the target data source, and the user behavior data generated by the target data source based on the question and answer data is obtained, including:
[0061] S120: When the target data source is a user interface or an application programming interface, the streaming question and answer data generated by the language model service cluster is forwarded to the corresponding user interface or application programming interface, and the streaming question and answer data is aggregated into corresponding complete question and answer data.
[0062] S121: When the target data source is an integrated development environment plug-in, the streaming question and answer data generated by the language model service cluster is forwarded to the corresponding integrated development environment plug-in, and the streaming question and answer data is aggregated into corresponding complete question and answer data.
[0063] S122: Obtain user behavior data generated by the integrated development environment plug-in based on the corresponding streaming question and answer data.
[0064] Specifically, when the target data source is a UI or API, the streaming question and answer data generated by the language model service cluster is forwarded to the corresponding UI or API, and the streaming question and answer data is aggregated into the corresponding complete question and answer data. When the target data source is an IDE plug-in, the streaming question and answer data generated by the LLM service cluster is forwarded to the corresponding IDE plug-in, and the streaming question and answer data is aggregated into the corresponding complete question and answer data. At the same time, the user behavior data generated by the IDE plug-in based on the corresponding streaming question and answer data is obtained.
[0065] It should be noted that the LLM service cluster uses streaming data (i.e., word-by-word generation and transmission) to output question and answer data mainly based on efficiency and experience optimization. Since it takes time to generate a complete answer, streaming transmission can gradually return the result, reduce the "blank period" of user waiting (such as chat dialogues displayed word by word), and enhance the sense of real-time interaction. At the same time, this method saves memory resources, avoids the overhead of one-time processing of long texts, and supports interruption in the middle (such as user stop request). Technically, the model generates content through word-by-word probability prediction, the server can push while generating, and the client renders synchronously, which is particularly suitable for network applications and real-time scenarios (such as translation, code completion). Therefore, in order to enable the target data source to obtain the generated answer in time and improve the efficiency of answer transmission, it is necessary to directly forward the streaming question and answer data to the target data source. At the same time, in order to better analyze the question and answer data output by the LLM service cluster, it is necessary to aggregate the streaming question and answer data into the corresponding complete question and answer data, so as to analyze the complete question and answer data. In this embodiment, there is no restriction on the specific process of aggregating the streaming question and answer data into the corresponding complete question and answer data.
[0066] Based on the above embodiments, in some embodiments, the streaming question and answer data is aggregated into corresponding complete question and answer data, including:
[0067] S123: Building a streaming data processing platform, and configuring data protocols, asynchronous processing, and concurrency parameters of the streaming data processing platform;
[0068] S124: receiving and parsing the streaming question-answering data generated by the language model service cluster through the streaming data processing platform;
[0069] S125: extracting target business fields in the streaming question-and-answer data by using a regular expression;
[0070] S126: splicing the target business fields transmitted in blocks in real time through a content splicing algorithm, and adding metadata information to each target business field to generate corresponding complete question and answer data;
[0071] The metadata information includes at least a timestamp, a block number and a data source.
[0072] like Figure 2As shown, in order to aggregate the streaming question and answer data into the corresponding complete question and answer data, a streaming data processing platform is pre-built in this embodiment to process the streaming data. It should be noted that the streaming data processing platform can choose to build a high-performance streaming processing framework, such as open source Apache Kafka Streams, Apache Flink or RedisStream and other streaming processing frameworks or lightweight streaming processing libraries. At the same time, the data protocol of the streaming data processing platform is configured, specifically based on the API provided by the streaming processing framework or library, to achieve the reception and parsing of the SSE protocol block-by-block data stream, and the asynchronous processing and concurrency parameters of the streaming processing framework or library, such as the thread pool size and concurrency, need to be configured to optimize data processing performance.
[0073] Furthermore, by receiving and parsing the streaming question and answer data generated by the LLM service cluster through the streaming data processing platform, and extracting the target business fields in the streaming question and answer data, such as response content, timestamp, user Internet Protocol address, session identifier, etc., through regular expressions, the data parsing efficiency and accuracy can be improved.
[0074] Finally, the target business fields transmitted in blocks are spliced in real time through the content splicing algorithm. In this embodiment, there is no restriction on the selected content splicing algorithm. For example, StringBuilder or StringBuffer is selected to splice the response content transmitted in blocks in real time. During the streaming process, metadata information is added to each target business field to generate the corresponding complete question and answer data. It should be noted that the metadata information includes at least the timestamp, block number and data source.
[0075] In addition, in order to ensure the reliability of the streaming data processing platform, after the initial construction is completed, the streaming data processing engine can be deployed in a cluster environment for performance testing and stress testing to ensure the engine's high throughput, low latency, and high reliability.
[0076] In this embodiment, a pre-built streaming data processing platform implements streaming data processing, and adopts a series of high-performance real-time parsing and asynchronous aggregation technologies such as SSE protocol block-by-block parsing, regular expression feature extraction, dynamic splicing of response content, block metadata marking, asynchronous processing and concurrency mechanism, to achieve low-latency processing of massive real-time data streams.
[0077] In order to monitor the user's code completion events and collect the user behavior data generated by the user based on the streaming question and answer data, based on the above embodiments, in some embodiments, obtaining the user behavior data generated by the integrated development environment plug-in based on the corresponding streaming question and answer data includes:
[0078] S130: Monitoring the user's code completion event through the integrated development environment plug-in.
[0079] S131: When the streaming question and answer data generated by the language model service cluster is received through the integrated development environment plug-in, a code completion event is determined to be triggered, and a code completion suggestion is determined according to the streaming question and answer data.
[0080] S132: Determine whether the user accepts the code completion suggestion within a preset period; if so, proceed to step S133; if not, proceed to step S134.
[0081] S133: Generate user behavior data representing the user's adoption of code completion suggestions.
[0082] S134: Generate user behavior data representing the user's rejection of the code completion suggestion.
[0083] like Figure 2 As shown, in this embodiment, the event monitoring interface provided by the IDE plug-in development framework is used to monitor the code completion events of the code editor through the behavior association collector, such as user input of keywords, triggering the display of code completion suggestions, etc. When the streaming question and answer data generated by the LLM service cluster is received through the IDE plug-in, the triggering code completion event is determined, and the code completion suggestion is determined according to the streaming question and answer data.
[0084] The behavior collection module of the behavior association collector further determines whether the user accepts the code completion suggestion within the preset period. If it is confirmed that the user accepts the code completion suggestion within the preset period, user behavior data representing the user's acceptance of the code completion suggestion is generated. If it is confirmed that the user does not accept the code completion suggestion within the preset period, user behavior data representing the user's rejection of the code completion suggestion is generated.
[0085] It should be noted that in this embodiment, there is no restriction on the preset period, for example, it can be set to 10s, that is, when the completion suggestion is displayed, a 10s timer is started. If the user does not adopt the suggestion within 10s (for example, does not select the suggestion or does not edit the code), it is determined that the user refuses to adopt the suggestion.
[0086] In summary, in this embodiment, through IDE plug-in integration and delay recording technology, accurate collection and delay determination of IDE code completion user behavior (explicit adoption and implicit rejection) are achieved.
[0087] In order to better store the question and answer data of the IDE plug-in and associate the question and answer data with the corresponding user behavior, based on the above embodiment, in some embodiments, after obtaining the user behavior data generated by the integrated development environment plug-in according to the corresponding streaming question and answer data, it also includes:
[0088] S135: Building a cache space and configuring the survival time of cache items in the cache space.
[0089] S136: storing the fingerprint mapping table based on the cache space.
[0090] The fingerprint mapping table contains multiple key-value pairs, where the key represents the request fingerprint and the value represents the complete question-and-answer data corresponding to the request fingerprint.
[0091] S137: Determine the request fingerprint of the user behavior data.
[0092] S138: Determine a key-value pair corresponding to the user behavior data in the fingerprint mapping table according to the request fingerprint of the user behavior data.
[0093] S139: Generate behavior question-answer associated data based on the user behavior data and the corresponding key-value pairs.
[0094] Specifically, build a behavior association module of the behavior association collector, select a high-performance memory database, such as Redis or Memcached to build a cache space, and configure the Least Recently Used (LRU) strategy of the cache space.
[0095] It should be noted that LRU is a common cache eviction algorithm used to manage data in the cache to optimize performance and resource utilization. The core idea of the LRU algorithm is: when the cache capacity reaches the upper limit, the data items that have been least recently accessed are removed first. In this way, the cache usually retains the data that has been frequently accessed recently, thereby improving the cache hit rate and system performance. In the LRU algorithm, each data item records the time or order of its last access. When data needs to be eliminated, the algorithm selects the data item that has not been accessed for the longest time to remove. The implementation of the LRU algorithm usually uses a combination of a bidirectional linked list and a hash table to achieve efficient access and update operations. This algorithm is widely used in various cache systems, such as virtual memory management of operating systems, database query caches, and page caches of Web applications. By using the LRU algorithm, the system can more effectively utilize limited cache resources and improve overall performance and response speed.
[0096] Specifically, the survival time of cache items in the cache space is configured based on the LRU strategy, for example, the survival time is set to 30 seconds. The LRU algorithm dynamically eliminates cold data, combined with a precise 30-second survival cycle, while ensuring data timeliness, maintaining a high cache hit rate and preventing memory overflow.
[0097] Furthermore, a fingerprint mapping table is stored based on the cache space. It is worth noting that the fingerprint mapping table is a mapping relationship table between request fingerprints and model call data, which is used to cache request context information in the recent period of time, and specifically includes multiple key-value pairs (key-value), where the key represents the request fingerprint and the value represents the complete question and answer data corresponding to the request fingerprint. In the specific implementation, a hash table can be used to construct the fingerprint mapping table.
[0098] Then, the request fingerprint of the user behavior data is determined, so as to determine the key-value pair corresponding to the user behavior data in the fingerprint mapping table according to the request fingerprint of the user behavior data. In this embodiment, there is no restriction on the method of determining the request fingerprint of the user behavior data, which depends on the specific implementation situation. Finally, after determining the key-value pair corresponding to the user behavior data, the behavior question-and-answer association data is generated according to the user behavior data and the corresponding key-value pair, so as to associate the question-and-answer data with the corresponding user behavior for subsequent analysis.
[0099] In this embodiment, by building a behavior association collector and proposing a request fingerprint-behavior association algorithm, an association is established between user behavior data and model call data, and a full-link user behavior model is constructed to facilitate subsequent complete analysis of the model call situation.
[0100] Based on the above embodiments, in some embodiments, determining the request fingerprint of user behavior data includes:
[0101] S140: When the streaming question and answer data generated by the language model service cluster is received, a corresponding question and answer timestamp is determined through an integrated development environment plug-in.
[0102] S141: Generate a request fingerprint of the user behavior data according to the generated user behavior data, the corresponding question and answer timestamp and a preset identifier generation algorithm.
[0103] In order to determine the request fingerprint of the user behavior data, in this embodiment, when the streaming question and answer data generated by the LLM service cluster is received and a round of dialogue is completed, the IDE plug-in will generate and determine the corresponding question and answer timestamp. Then the IDE plug-in will generate an algorithm based on the generated user behavior data, the corresponding question and answer timestamp and the preset identifier to generate the request fingerprint of the user behavior data.
[0104] It should be noted that the preset identifier generation algorithm here should be the same as the preset identifier generation algorithm for generating the request fingerprint of the model call request, so as to ensure that the same request fingerprint is generated to facilitate associating the question and answer data with the corresponding user behavior.
[0105] Before performing full-link multi-dimensional analysis based on question-and-answer data and corresponding user behavior data, after obtaining the user behavior data generated by the target data source based on the question-and-answer data, it also includes:
[0106] S142: When the target data source is a user interface or an application programming interface, the corresponding complete question and answer data is stored in a pre-built storage space.
[0107] S143: When the target data source is an integrated development environment plug-in, the corresponding behavior question and answer associated data is stored in a pre-built storage space.
[0108] In order to better save the complete question and answer data / behavior question and answer related data, in this embodiment, when the target data source is UI or API, the corresponding complete question and answer data is stored in the pre-built storage space. When the target data source is an IDE plug-in, the corresponding behavior question and answer related data is stored in the pre-built storage space. The storage process is described in detail below:
[0109] In some embodiments, storing the complete question and answer data / behavior question and answer associated data in a storage space includes:
[0110] S144: Build hot data storage space, warm data storage space, and cold data storage space.
[0111] S145: Storing the complete question and answer data / behavior question and answer associated data in the hot data storage space.
[0112] S146: Monitor access information of complete question and answer data / behavioral question and answer associated data.
[0113] S147: Downgrade the storage of the complete question and answer data / behavior question and answer associated data to a warm data storage space or a cold data storage space according to the access information.
[0114] like Figure 2 As shown, in the specific implementation, an intelligent storage controller is deployed to build hot data storage space, warm data storage space and cold data storage space respectively. In this embodiment, there is no restriction on the storage medium selected for each storage space. For example, Redis Cluster or Memcached Cluster can be selected as hot data storage space, Apache Kafka or RabbitMQ can be selected as warm data storage space, and ClickHouse Cluster or HDFS Cluster can be selected as cold data storage space.
[0115] Further configure the hierarchical storage strategy. Specifically, store the complete question and answer data / behavioral question and answer associated data in the hot data storage space, and then monitor the access information of the complete question and answer data / behavioral question and answer associated data. It should be noted that the access information at least includes the access frequency. Finally, the complete question and answer data / behavioral question and answer associated data is downgraded and stored in the warm data storage space or the cold data storage space according to the access information. It should be noted that in this embodiment, there is no restriction on the specific process of downgrading the complete question and answer data / behavioral question and answer associated data to the warm data storage space or the cold data storage space according to the access information, which depends on the specific implementation situation.
[0116] In addition, in order to ensure the availability of the intelligent storage controller, the intelligent storage controller can also be deployed in a cluster environment for performance testing, stress testing, and disaster recovery testing to ensure the high performance, high reliability, high scalability, and data security of the storage controller.
[0117] In this embodiment, by deploying an intelligent storage controller, using hierarchical storage media to build a data hierarchical storage system, and creating a dynamic degradation mechanism, efficient data storage management is guaranteed.
[0118] Based on the above embodiments, in some embodiments, the complete question and answer data / behavior question and answer associated data is downgraded and stored in a warm data storage space or a cold data storage space according to the access information, including:
[0119] S148: Determine whether there is access to the complete question and answer data / behavioral question and answer associated data within the first preset time based on the access information; if so, return to step S148; if not, proceed to step S149.
[0120] S149: Downgrade the complete question and answer data / behavioral question and answer associated data to the warm data storage space, and determine whether there is access to the complete question and answer data / behavioral question and answer associated data within the second preset time based on the access information; if so, return to step S149; if not, enter step S150.
[0121] S150: Degrade the storage of the complete question and answer data / behavior question and answer associated data to a cold data storage space.
[0122] The first preset time is shorter than the second preset time.
[0123] In order to achieve downgraded storage of complete question and answer data / behavior question and answer associated data, in this embodiment, it is specifically determined based on the access information whether there is access to the complete question and answer data / behavior question and answer associated data within the first preset time.
[0124] If it is confirmed that there is access to the complete question and answer data / behavioral question and answer associated data within the first preset time, the complete question and answer data / behavioral question and answer associated data is considered to be high-frequency access data and still needs to be saved in the hot data storage space, and the process returns to the step of determining whether there is access to the complete question and answer data / behavioral question and answer associated data within the first preset time based on the access information, so as to continue monitoring the access information of the data. If it is confirmed that there is no access to the complete question and answer data / behavioral question and answer associated data within the first preset time, the access frequency of the complete question and answer data / behavioral question and answer associated data is considered to be average. In order to optimize data storage, the complete question and answer data / behavioral question and answer associated data needs to be downgraded and stored in the warm data storage space, and based on the access information, it is determined whether there is access to the complete question and answer data / behavioral question and answer associated data within the second preset time.
[0125] If it is confirmed that there is access to the complete question and answer data / behavioral question and answer associated data within the second preset time, it is considered that the access frequency of the complete question and answer data / behavioral question and answer associated data is general, and it still needs to be saved in the warm data storage space, and return to the step of determining whether there is access to the complete question and answer data / behavioral question and answer associated data within the second preset time based on the access information, so as to continue monitoring the access information of the data. If it is confirmed that there is no access to the complete question and answer data / behavioral question and answer associated data within the second preset time, it is considered that the access frequency of the complete question and answer data / behavioral question and answer associated data is low. In order to optimize data storage, the complete question and answer data / behavioral question and answer associated data needs to be downgraded to cold data storage space. In this way, downgraded storage of data is achieved.
[0126] It should be noted that in this embodiment, no restrictions are imposed on the first preset time and the second preset time. It is only necessary to ensure that the first preset time is less than the second preset time.
[0127] On the basis of the above embodiments, in some embodiments, full-link multi-dimensional analysis is performed based on question-answer data and corresponding user behavior data, including:
[0128] S151: Build a data analysis platform and corresponding data warehouse.
[0129] S152: Obtain the complete question and answer data / behavior question and answer associated data in the storage space and store them in the data warehouse.
[0130] S153: Perform multi-dimensional analysis on the complete question and answer data / behavioral question and answer related data through the data analysis platform to generate call analysis results of the language model service cluster.
[0131] like Figure 2 As shown, in order to implement LLM call data analysis, a data analysis platform and a corresponding data warehouse are specifically built. For example, ClickHouse Cluster or Hadoop+Spark / Flink is selected as the data warehouse of the data analysis platform, and other frameworks can also be used, which is not limited in this embodiment.
[0132] Subsequently, the data analysis platform obtains the complete question and answer data / behavior question and answer related data in the storage space and stores it in the data warehouse. Finally, a multi-dimensional analysis is performed on the complete question and answer data / behavior question and answer related data to generate the call analysis results of the language model service cluster. The following is a detailed description of the data storage and analysis process:
[0133] In some embodiments, obtaining complete question-and-answer data / behavior question-and-answer associated data in a storage space and storing them in a data warehouse includes:
[0134] S154: Perform data cleaning and standardization on the complete question and answer data / behavioral question and answer associated data through the data analysis platform to obtain pre-processed complete question and answer data / behavioral question and answer associated data.
[0135] S155: Storing the pre-processed complete question and answer data / behavior question and answer associated data in the data warehouse.
[0136] Specifically, based on big data processing frameworks such as Spark and Flink, the data analysis platform can clean, convert and standardize the complete question and answer data / behavior question and answer related data, such as data format conversion, outlier processing, data desensitization, data normalization, and efficiently store the cleaned structured data in the data warehouse.
[0137] Correspondingly, the data analysis platform conducts multi-dimensional analysis on the complete question and answer data / behavioral question and answer related data, including:
[0138] S156: Perform multi-dimensional analysis on the complete question and answer data / behavioral question and answer associated data through a variety of data processing and analysis algorithms to generate call analysis results.
[0139] The data analysis platform uses a variety of data processing and analysis algorithms to conduct multi-dimensional analysis of complete question-and-answer data / behavioral question-and-answer related data. Specific analysis methods include but are not limited to call volume statistical analysis, user behavior analysis, model performance evaluation, real-time monitoring and early warning, A / B testing and effect evaluation, error and anomaly analysis, resource consumption analysis, security and compliance audit, etc., and finally generates the call analysis results of the model. The following is an explanation of the analysis process:
[0140] First, for the LLM service cluster, the call volume statistical analysis can reveal the frequency of model requests, peak usage periods, and the popularity of different functional modules. These data help optimize model deployment, ensure stable service during high demand, and guide future model expansion and resource planning. Secondly, by analyzing the interaction between users and LLM, including query types, feedback, and usage habits, we can gain a deep understanding of user needs and preferences, which can be used to improve the response quality of the model, optimize the user interface, and develop new features that better meet user expectations. On the other hand, the performance evaluation of the model and its accuracy, response time, relevance and fluency of the generated text; by regularly evaluating and comparing models of different versions or configurations, we can identify areas for improvement and ensure that the model continues to provide high-quality output. Subsequently, real-time monitoring of the operating status of LLM, including resource usage, response delay, and error rate, helps to promptly discover and resolve potential problems; at the same time, by setting up an early warning mechanism, we can respond quickly when performance degrades or failures occur, ensuring the continuity and reliability of services. In addition, through A / B testing, we can compare the impact of different versions of LLM or specific functions on user experience and business indicators, which helps to determine the best model configuration or function design, optimize user satisfaction, and improve the value of the model in practical applications. Furthermore, by collecting and analyzing the errors and anomalies generated by LLM during operation, common problem patterns and potential defects can be identified; by fixing these problems, the stability and robustness of the model can be improved, and the errors and interruptions encountered by users can be reduced. Monitoring the resource consumption of LLM, such as computing resources, storage, and bandwidth usage, helps to optimize resource allocation and cost management; by analyzing resource usage patterns, optimization opportunities can be identified, resource utilization can be improved, and the model can be ensured to run efficiently within the budget. Finally, security and compliance audits are conducted on LLM to ensure that it complies with relevant regulations and standards when processing user data and generating content. This includes reviewing data privacy protection measures, content filtering mechanisms, and access control policies to protect user rights and avoid legal risks. Through continuous auditing and improvement, user trust can be established and the model can be ensured to run in a secure and compliant environment.
[0141] In this embodiment, the data analysis platform performs data reception, cleaning and storage, call volume statistical analysis, user behavior analysis, model performance evaluation, real-time monitoring and early warning, A / B testing and effect evaluation, error and exception analysis, resource consumption analysis, and security and compliance auditing on complete question and answer data / behavioral question and answer related data, providing multi-dimensional, intelligent, and comprehensive data analysis capabilities.
[0142] In addition, in order to ensure the reliability of the data analysis platform, it can also be deployed in a cluster environment for functional testing, performance testing, stress testing, and stability testing to ensure the platform's data processing and analysis capabilities, multi-dimensional intelligent analysis functions, and real-time monitoring and early warning capabilities, with high concurrency, high reliability, and high availability.
[0143] In order to enable the user to view the call analysis results of the language model more intuitively, based on the above embodiments, in some embodiments, after the call analysis results of the language model service cluster are generated, the following is further included:
[0144] S157: Develop a data visualization display interface for the data analysis platform through data visualization tools.
[0145] S158: Output the call analysis results through the data visualization display interface.
[0146] Specifically, based on the data analysis platform, we select open source data visualization tools to develop diversified, custom visualization reports, dashboards, monitoring screens and other data visualization display interfaces to present data analysis results in real time. For example, call volume trend charts, user behavior distribution charts, model performance indicator dashboards, real-time monitoring dashboards, A / B test comparison reports, etc., support custom visualization reports and dashboards, and users can flexibly configure and display key indicators according to their needs. Lower the threshold for data understanding and use, so that users can intuitively and clearly understand the application status, performance and potential problems of LLM, and provide data support for decision-making.
[0147] In addition, in order to facilitate users to query, analyze and use the collected LLM call data and analysis results, the method also includes:
[0148] S159: Based on the network application programming interface framework, develop a data analysis interface for the data analysis platform and open the data analysis function of the data analysis platform.
[0149] S160: Develop a visualization interface based on the front-end framework to facilitate users to directly access the call analysis results in the data analysis platform through the visualization interface and data analysis interface.
[0150] Specifically, based on open source web application programming interface frameworks, such as FastAPI, Spring Boot, Django REST framework and other API frameworks, data analysis interfaces (RESTful APIs) are developed for data analysis platforms, and data analysis functions of data analysis platforms are opened. Multi-dimensional data filtering, sorting and paging functions are provided, such as call volume query API, user behavior analysis API, model performance evaluation API, real-time monitoring data API, A / B test result API, etc. It is convenient for users to obtain the required data on demand, and supports SDKs of multiple programming languages and development platforms, lowering the threshold for API use, while supporting JSON data format and multiple programming languages.
[0151] Furthermore, a visualization interface is developed based on the front-end framework so that users can directly access the call analysis results in the data analysis platform through the visualization interface and data analysis interface. For example, a front-end framework such as React, Vue, and Angular is selected to develop a Web-based visualization interface, so that users can directly access the data analysis platform through a browser, interactively query, analyze, and visualize LLM call data, customize visualization reports and dashboards, and export data analysis reports.
[0152] In addition, in order to ensure the availability of the data analysis interface, the data analysis interface can be deployed in a Web server cluster for functional testing, performance testing, and security testing to ensure the interface's ease of use, high performance, high reliability, and high security.
[0153] In this embodiment, by providing a standardized API interface and a visual interface, it is convenient for users to query, analyze and use the collected LLM call data and analysis results, thereby realizing open sharing and application of data.
[0154] In order to make the LLM service cluster run better, in some embodiments, after generating the call analysis result of the language model service cluster, the method further includes:
[0155] S161: Obtain load information of the language model service cluster;
[0156] S162: Generate a distribution strategy for language model call requests according to the call analysis result and the load information, so as to process new language model call requests based on the distribution strategy.
[0157] Specifically, the load information of the language model service cluster is obtained, including but not limited to processor usage, memory usage, disk input and output, network traffic, etc. Subsequently, the request allocation strategy is dynamically adjusted according to the call analysis results and load information, so as to process new language model call requests based on the allocation strategy.
[0158] In addition, in some embodiments, it also includes:
[0159] S163: Monitor the health status of the language model service cluster.
[0160] S164: When a faulty server is detected in the language model service cluster, the language model call request being processed in the faulty server is transferred to the remaining healthy servers for execution.
[0161] Specifically, the health status of the LLM service cluster is checked regularly to promptly detect and handle faulty servers to ensure high availability of the system. When a server in the LLM service cluster fails, the language model call requests being processed by the failed server are transferred to the remaining healthy servers for execution, thereby avoiding interruption of request processing.
[0162] In summary, performing load balancing and fault handling on the LLM service cluster based on the call analysis results can effectively improve the performance, stability, and resource utilization of LLM, thereby better meeting user needs and business requirements.
[0163] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0164] Figure 3 Schematic diagram of an information collection and analysis device provided by an embodiment of the present invention. Figure 3 As shown, the device comprises:
[0165] The monitoring module 10 is used to monitor language model call requests transmitted by multiple data sources through a gateway with multi-protocol identification function and a dynamic traffic processing platform; wherein the data source includes at least a user interface, an application programming interface and an integrated development environment plug-in.
[0166] The forwarding module 11 is used to forward the language model call request to the language model service cluster when receiving the language model call request transmitted by the target data source, so as to generate corresponding question and answer data through the language model service cluster.
[0167] The acquisition module 12 is used to forward the question and answer data to the target data source and obtain the user behavior data generated by the target data source based on the question and answer data.
[0168] The analysis module 13 is used to perform full-link multi-dimensional analysis based on the question-answer data and the corresponding user behavior data to generate a call analysis result of the language model service cluster.
[0169] In some embodiments, the monitoring module 10 includes:
[0170] A first monitoring submodule is used to monitor the flow data transmitted by each data source through the corresponding protocol based on the gateway and the dynamic flow processing platform;
[0171] The first identification submodule is used to identify language model calling requests that meet language model calling characteristics in each flow data through deep packet inspection technology.
[0172] In some embodiments, the forwarding module 11 includes:
[0173] A first generation submodule is used to generate a request fingerprint of a language model call request by using a preset identifier generation algorithm; wherein the request fingerprint includes at least metadata information of a user Internet Protocol address, a client type, an entry type, a model version, and a request timestamp;
[0174] The first adding submodule is used to add the request fingerprint to the request header of the language model call request;
[0175] The first forwarding submodule is used to forward the language model call request to the language model service cluster according to the request fingerprint in the corresponding request header.
[0176] In some embodiments, the first forwarding submodule includes:
[0177] The first configuration submodule is used to configure the proxy delivery instructions between the gateway and the dynamic traffic processing platform;
[0178] The second forwarding submodule is used to control the gateway and the dynamic traffic processing platform to forward the language model call request to the language model service cluster according to the request fingerprint, entry type and model version based on the proxy delivery instruction.
[0179] In some embodiments, the acquisition module 12 includes:
[0180] A first aggregation submodule, configured to forward the streaming question and answer data generated by the language model service cluster to the corresponding user interface or application programming interface when the target data source is a user interface or an application programming interface, and aggregate the streaming question and answer data into corresponding complete question and answer data;
[0181] A second aggregation submodule is used for forwarding the streaming question and answer data generated by the language model service cluster to the corresponding integrated development environment plug-in when the target data source is an integrated development environment plug-in, and aggregating the streaming question and answer data into corresponding complete question and answer data;
[0182] The first acquisition submodule is used to acquire user behavior data generated by the integrated development environment plug-in according to the corresponding streaming question and answer data.
[0183] In some embodiments, the first aggregation submodule and the second aggregation submodule include:
[0184] The first building submodule is used to build a streaming data processing platform and configure the data protocol, asynchronous processing and concurrency parameters of the streaming data processing platform;
[0185] A first parsing submodule, configured to receive and parse the streaming question-answering data generated by the language model service cluster through the streaming data processing platform;
[0186] A first extraction submodule, configured to extract target business fields in the streaming question-and-answer data by using a regular expression;
[0187] The splicing submodule is used to splice the target business fields transmitted in blocks in real time through the content splicing algorithm, and add metadata information to each target business field to generate corresponding complete question and answer data;
[0188] The metadata information includes at least a timestamp, a block number and a data source.
[0189] In some embodiments, the first acquisition submodule includes:
[0190] The first monitoring submodule is used to monitor the user's code completion events through the integrated development environment plug-in;
[0191] A first determination submodule is used to determine a triggering code completion event when receiving streaming question and answer data generated by a language model service cluster through an integrated development environment plug-in, and to determine a code completion suggestion according to the streaming question and answer data;
[0192] The first judgment submodule is used to judge whether the user accepts the code completion suggestion within a preset period; if so, generate user behavior data representing the user's adoption of the code completion suggestion; if not, generate user behavior data representing the user's rejection of the code completion suggestion.
[0193] In some embodiments, it also includes:
[0194] The second submodule is used to build a cache space and configure the survival time of cache items in the cache space;
[0195] A first storage submodule is used to store a fingerprint mapping table based on a cache space; wherein the fingerprint mapping table contains a plurality of key-value pairs, the key represents the requested fingerprint, and the value represents the complete question-answer data corresponding to the requested fingerprint;
[0196] A second determination submodule is used to determine the request fingerprint of the user behavior data;
[0197] A third determination submodule is used to determine a key-value pair corresponding to the user behavior data in the fingerprint mapping table according to the request fingerprint of the user behavior data;
[0198] Generate behavioral question-answer related data based on user behavior data and corresponding key-value pairs.
[0199] In some embodiments, the second determining submodule includes:
[0200] A fourth determination submodule is used to determine a corresponding question and answer timestamp through an integrated development environment plug-in when receiving streaming question and answer data generated by the language model service cluster;
[0201] The second generating submodule is used to generate a request fingerprint of the user behavior data according to the generated user behavior data, the corresponding question and answer timestamp and a preset identifier generation algorithm.
[0202] In some embodiments, it also includes:
[0203] A second storage submodule, for storing the corresponding complete question and answer data in a pre-built storage space when the target data source is a user interface or an application programming interface;
[0204] The third storage submodule is used to store the corresponding behavior question and answer associated data in a pre-built storage space when the target data source is an integrated development environment plug-in.
[0205] In some embodiments, the second storage submodule and the third storage submodule include:
[0206] The third submodule is used to build hot data storage space, warm data storage space and cold data storage space;
[0207] A fourth storage submodule is used to store the complete question and answer data / behavior question and answer associated data into a hot data storage space;
[0208] The second monitoring submodule is used to monitor access information of complete question and answer data / behavior question and answer associated data;
[0209] The downgraded storage submodule is used to downgrade the storage of complete question and answer data / behavior question and answer associated data to warm data storage space or cold data storage space according to the access information.
[0210] In some embodiments, the downgrading storage submodule includes:
[0211] The second judgment submodule is used to judge whether there is access to the complete question and answer data / behavior question and answer associated data within the first preset time according to the access information; if so, the second judgment submodule is triggered; if not, the third judgment submodule is triggered;
[0212] The third judgment submodule is used to downgrade the complete question and answer data / behavioral question and answer associated data to the warm data storage space, and judge whether there is access to the complete question and answer data / behavioral question and answer associated data within the second preset time according to the access information; if so, the third judgment submodule is triggered; if not, the complete question and answer data / behavioral question and answer associated data is downgraded and stored in the cold data storage space.
[0213] The first preset time is shorter than the second preset time.
[0214] In some embodiments, the analysis module 13 includes:
[0215] The fourth sub-module is used to build a data analysis platform and a corresponding data warehouse;
[0216] The second acquisition submodule is used to acquire the complete question and answer data / behavior question and answer associated data in the storage space and store it in the data warehouse;
[0217] The first analysis submodule is used to perform multi-dimensional analysis on the complete question and answer data / behavioral question and answer related data through the data analysis platform to generate a call analysis result of the language model service cluster.
[0218] In some embodiments, the second acquisition submodule includes:
[0219] The data preprocessing submodule is used to clean and standardize the complete question and answer data / behavior question and answer related data through the data analysis platform to obtain the preprocessed complete question and answer data / behavior question and answer related data;
[0220] A fifth storage submodule is used to store the pre-processed complete question and answer data / behavior question and answer associated data into a data warehouse;
[0221] Correspondingly, the first analysis submodule includes:
[0222] The multi-dimensional analysis submodule is used to perform multi-dimensional analysis on the complete question and answer data / behavioral question and answer associated data through a variety of data processing and analysis algorithms to generate call analysis results.
[0223] In some embodiments, the apparatus further comprises:
[0224] The first development submodule is used to develop a data visualization display interface for the data analysis platform through a data visualization tool;
[0225] The first output submodule is used to output the call analysis results through a data visualization display interface.
[0226] In some embodiments, the apparatus further comprises:
[0227] The second development submodule is used to develop a data analysis interface for the data analysis platform based on the network application programming interface framework, and to open up the data analysis function of the data analysis platform;
[0228] The third development submodule is used to develop a visualization interface based on the front-end framework, so that users can directly access the call analysis results in the data analysis platform through the visualization interface and the data analysis interface.
[0229] In some embodiments, the apparatus further comprises:
[0230] The third acquisition submodule is used to obtain the load information of the language model service cluster;
[0231] The third generation submodule is used to generate an allocation strategy for language model call requests according to the call analysis results and load information, so as to process new language model call requests based on the allocation strategy.
[0232] In some embodiments, the apparatus further comprises:
[0233] The third monitoring submodule is used to monitor the health status of the language model service cluster;
[0234] The transfer submodule is used to transfer the language model call request being processed in the faulty server to the remaining healthy servers for execution when a faulty server is detected in the language model service cluster.
[0235] It can be understood that the description of the features in the embodiment corresponding to the information collection and analysis device can refer to the relevant description of the embodiment corresponding to the information collection and analysis method, and will not be repeated here.
[0236] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned information collection and analysis method embodiments.
[0237] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned information collection and analysis method embodiments when running.
[0238] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0239] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned information collection and analysis method embodiments are implemented.
[0240] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned information collection and analysis method embodiments are implemented.
[0241] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0242] The above is a detailed introduction to the information collection and analysis method, device, equipment, medium and product provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. An information collection and analysis method, characterized in that: include: Monitoring language model call requests transmitted by multiple data sources through a gateway and a dynamic traffic processing platform with multi-protocol identification capabilities; wherein the data sources include at least a user interface, an application programming interface, and an integrated development environment plug-in; When receiving a language model call request transmitted by a target data source, forwarding the language model call request to a language model service cluster to generate corresponding question and answer data through the language model service cluster; Forwarding the question and answer data to the target data source, and obtaining user behavior data generated by the target data source based on the question and answer data; A full-link multi-dimensional analysis is performed based on the question and answer data and the corresponding user behavior data to generate a call analysis result of the language model service cluster.
2. The information collection and analysis method according to claim 1, characterized in that: Through a gateway with multi-protocol identification capabilities and a dynamic traffic processing platform, language model call requests transmitted by multiple data sources are monitored, including: Based on the gateway and the dynamic traffic processing platform, the traffic data transmitted by each data source through the corresponding protocol is monitored; The language model calling request that meets the language model calling characteristics in each of the traffic data is identified through deep message inspection technology.
3. The information collection and analysis method according to claim 2, characterized in that: When a language model call request transmitted by a target data source is received, the language model call request is forwarded to a language model service cluster, including: Generate a request fingerprint of the language model call request by a preset identifier generation algorithm; wherein the request fingerprint includes at least metadata information of a user Internet Protocol address, a client type, an entry type, a model version, and a request timestamp; Adding the request fingerprint to the request header of the language model call request; According to the request fingerprint in the corresponding request header, the language model call request is forwarded to the language model service cluster.
4. The information collection and analysis method according to claim 3 is characterized in that: Forwarding the language model call request to the language model service cluster according to the request fingerprint in the corresponding request header, including: Configuring proxy delivery instructions between the gateway and the dynamic traffic processing platform; Based on the proxy delivery instruction, the gateway and the dynamic traffic processing platform are controlled to forward the language model call request to the language model service cluster according to the request fingerprint, the entry type and the model version.
5. The information collection and analysis method according to claim 1, characterized in that: Forwarding the question and answer data to the target data source, and obtaining user behavior data generated by the target data source based on the question and answer data, including: When the target data source is a user interface or an application programming interface, forwarding the streaming question and answer data generated by the language model service cluster to the corresponding user interface or application programming interface, and aggregating the streaming question and answer data into corresponding complete question and answer data; When the target data source is an integrated development environment plug-in, forwarding the streaming question and answer data generated by the language model service cluster to the corresponding integrated development environment plug-in, and aggregating the streaming question and answer data into corresponding complete question and answer data; The user behavior data generated by the integrated development environment plug-in according to the corresponding streaming question and answer data is obtained.
6. The information collection and analysis method according to claim 5, characterized in that: Aggregate streaming question-answering data into corresponding complete question-answering data, including: Building a streaming data processing platform, and configuring data protocols, asynchronous processing, and concurrency parameters of the streaming data processing platform; Receiving and parsing the streaming question-answering data generated by the language model service cluster through the streaming data processing platform; Extract target business fields from streaming question-and-answer data using regular expressions; The target business fields transmitted in blocks are spliced in real time by a content splicing algorithm, and metadata information is added to each target business field to generate the corresponding complete question and answer data; The metadata information at least includes a timestamp, a block number and a data source.
7. The information collection and analysis method according to claim 5, characterized in that: Acquiring the user behavior data generated by the integrated development environment plug-in according to the corresponding streaming question-and-answer data, including: Monitoring the user's code completion events through the integrated development environment plug-in; When the streaming question and answer data generated by the language model service cluster is received through the integrated development environment plug-in, a code completion event is determined to be triggered, and a code completion suggestion is determined according to the streaming question and answer data; Determining whether the user accepts the code completion suggestion within a preset period; If yes, generating the user behavior data representing that the user adopts the code completion suggestion; If not, the user behavior data representing that the user rejects the code completion suggestion is generated.
8. The information collection and analysis method according to claim 7, characterized in that: After obtaining the user behavior data generated by the integrated development environment plug-in according to the corresponding streaming question and answer data, the method further includes: Building a cache space and configuring the survival time of cache items in the cache space; A fingerprint mapping table is stored based on the cache space; wherein the fingerprint mapping table contains a plurality of key-value pairs, the key represents the request fingerprint, and the value represents the complete question-and-answer data corresponding to the request fingerprint; Determining a request fingerprint of the user behavior data; Determine, according to the request fingerprint of the user behavior data, a key-value pair corresponding to the user behavior data in the fingerprint mapping table; Behavior question-answer associated data is generated according to the user behavior data and the corresponding key-value pairs.
9. The information collection and analysis method according to claim 8, characterized in that: Determining the request fingerprint of the user behavior data includes: When receiving the streaming question and answer data generated by the language model service cluster, determining the corresponding question and answer timestamp through the integrated development environment plug-in; A request fingerprint of the user behavior data is generated according to the generated user behavior data, the corresponding question and answer timestamp and a preset identifier generation algorithm.
10. The information collection and analysis method according to claim 8, characterized in that: Before performing full-link multi-dimensional analysis based on the question-and-answer data and the corresponding user behavior data, after obtaining the user behavior data generated by the target data source based on the question-and-answer data, the method further includes: When the target data source is a user interface or an application programming interface, storing the corresponding complete question and answer data in a pre-built storage space; When the target data source is an integrated development environment plug-in, the corresponding behavior question and answer associated data is stored in the pre-constructed storage space.
11. The information collection and analysis method according to claim 10, characterized in that: Storing the complete question and answer data / the behavior question and answer associated data into the storage space includes: Build hot data storage space, warm data storage space and cold data storage space; storing the complete question and answer data / the behavioral question and answer associated data in the hot data storage space; Monitoring access information of the complete question and answer data / the behavioral question and answer associated data; The complete question and answer data / the behavior question and answer associated data are downgraded and stored in the warm data storage space or the cold data storage space according to the access information.
12. The information collection and analysis method according to claim 11, characterized in that: Degrading and storing the complete question and answer data / the behavior question and answer associated data to the warm data storage space or the cold data storage space according to the access information includes: Determining, according to the access information, whether there is access to the complete question and answer data / the behavioral question and answer associated data within a first preset time; If it is confirmed that there is access to the complete question and answer data / the behavior question and answer associated data within the first preset time, return to the step of determining whether there is access to the complete question and answer data / the behavior question and answer associated data within the first preset time according to the access information; If it is confirmed that there is no access to the complete question and answer data / the behavioral question and answer associated data within the first preset time, the complete question and answer data / the behavioral question and answer associated data are downgraded and stored in the warm data storage space, and it is determined based on the access information whether there is access to the complete question and answer data / the behavioral question and answer associated data within the second preset time; If it is confirmed that there is access to the complete question and answer data / the behavior question and answer associated data within the second preset time, return to the step of determining whether there is access to the complete question and answer data / the behavior question and answer associated data within the second preset time according to the access information; If it is confirmed that there is no access to the complete question and answer data / the behavior question and answer associated data within the second preset time, downgrading the complete question and answer data / the behavior question and answer associated data to the cold data storage space; Wherein, the first preset time is shorter than the second preset time.
13. The information collection and analysis method according to claim 10, characterized in that: Performing full-link multi-dimensional analysis based on the question-and-answer data and the corresponding user behavior data, including: Build a data analysis platform and corresponding data warehouse; Acquire the complete question and answer data / the behavioral question and answer associated data in the storage space, and store them in the data warehouse; The data analysis platform performs a multi-dimensional analysis on the complete question and answer data / the behavioral question and answer associated data to generate the call analysis result of the language model service cluster.
14. The information collection and analysis method according to claim 13, characterized in that: Acquiring the complete question and answer data / the behavior question and answer associated data in the storage space and storing them in the data warehouse includes: Performing data cleaning and standardization processing on the complete question and answer data / the behavioral question and answer associated data through the data analysis platform to obtain the pre-processed complete question and answer data / the behavioral question and answer associated data; storing the preprocessed complete question and answer data / the behavioral question and answer associated data in the data warehouse; Correspondingly, the data analysis platform performs a multi-dimensional analysis on the complete question and answer data / the behavioral question and answer associated data, including: The complete question and answer data / the behavioral question and answer associated data are subjected to multi-dimensional analysis through a variety of data processing and analysis algorithms to generate the call analysis result.
15. The information collection and analysis method according to claim 14, characterized in that: After generating the call analysis result of the language model service cluster, the method further includes: Developing a data visualization display interface for the data analysis platform through a data visualization tool; The call analysis result is output through the data visualization display interface.
16. The information collection and analysis method according to claim 15, characterized in that: Also includes: Based on the network application programming interface framework, develop a data analysis interface for the data analysis platform and open up the data analysis function of the data analysis platform; A visualization interface is developed based on the front-end framework so that users can directly access the call analysis results in the data analysis platform through the visualization interface and the data analysis interface.
17. The information collection and analysis method according to any one of claims 1 to 16, characterized in that: After generating the call analysis result of the language model service cluster, the method further includes: Obtaining load information of the language model service cluster; A distribution strategy for language model call requests is generated according to the call analysis result and the load information, so as to process new language model call requests based on the distribution strategy.
18. The information collection and analysis method according to claim 17, characterized in that: Also includes: Monitoring the health status of the language model service cluster; When it is detected that there is a faulty server in the language model service cluster, the language model call request being processed in the faulty server is transferred to the remaining healthy servers for execution.
19. An information collection and analysis device, characterized in that: include: A monitoring module, used to monitor language model call requests transmitted by multiple data sources through a gateway with multi-protocol identification function and a dynamic traffic processing platform; wherein the data source includes at least a user interface, an application programming interface, and an integrated development environment plug-in; A forwarding module, configured to forward, when receiving a language model call request transmitted by a target data source, the language model call request to the language model service cluster, so as to generate corresponding question and answer data through the language model service cluster; An acquisition module, used to forward the question and answer data to the target data source, and acquire user behavior data generated by the target data source based on the question and answer data; An analysis module is used to perform full-link multi-dimensional analysis based on the question and answer data and the corresponding user behavior data to generate a call analysis result of the language model service cluster.
20. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the information collection and analysis method as described in any one of claims 1 to 18 when executing the computer program.
21. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the information collection and analysis method according to any one of claims 1 to 18.
22. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the information collection and analysis method according to any one of claims 1 to 18 are implemented.
Citation Information
Patent Citations
Evaluation method and device for distributed system, electronic equipment and storage medium
CN114124759A
Hospital quality monitoring data analysis and fine management system and method
CN117038025A
Large-scale artificial intelligence language model interface calling usage monitoring and analyzing platform
CN117851151A
Lake-bin-chain integrated efficient credible big data storage and analysis system
CN118838551A
BOM data multi-dimensional analysis system and analysis method based on large language model
CN119004372A
Cited By
Rapid data acquisition method based on fusion of intelligent number asking and industrial connection technology
CN120358288A
A fast data acquisition method based on the integration of intelligent data processing and industrial connection technology
CN120358288B
Question and answer pair generation method and device for function codes
CN120470099A
Service performance evaluation method, computer program product and electronic device
CN120723342A
Large language model dynamic adaptation method and system based on Java
CN121012756A