Intelligent dialogue methods, devices, equipment, and storage media based on streaming AI
By identifying and hierarchically caching key data units in a streaming AI dialogue system, and combining interruption location and context state information, the system enables breakpoint resumption of dialogue content, solving the problem of dialogue loss caused by network interruption and improving user experience and system efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN WEIGEYUN TECH CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-06-02
AI Technical Summary
Existing streaming AI dialogue systems cannot effectively resume the dialogue process after a network interruption, resulting in the loss of dialogue content, which affects user experience and wastes resources.
The key data unit set is determined based on the correlation between data units in the data stream, and a storage strategy identifier is assigned to it. The data units are then stored in a multi-level cache space. When the network is interrupted, the interruption location and context state information are obtained and stored in association. After the network connection is restored, the key and historical data units are obtained from the multi-level cache, the data stream is reconstructed, and the data units are merged.
It enables breakpoint resumption of streaming AI data streams, avoiding the need to regenerate all content, saving computing resources and network bandwidth, reducing user waiting time, and improving user experience and system response efficiency.
Smart Images

Figure CN122137918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an intelligent dialogue method, apparatus, device, and storage medium based on streaming AI. Background Technology
[0002] With the rapid development of artificial intelligence technology, streaming AI dialogue systems have been widely used in fields such as intelligent customer service, virtual assistants, and online education. Streaming AI dialogue systems generate responses in real-time, word-by-word output, providing users with an instant interactive feedback experience. However, the streaming AI dialogue process is highly dependent on the stability of the network connection, and in practical applications, it often faces problems such as network fluctuations and connection interruptions.
[0003] Existing streaming AI dialogue systems cannot effectively resume dialogue after network interruptions. When the network connection is interrupted, the ongoing dialogue content is completely lost, and the user needs to re-initiate the request. This results in the inability to reuse the generated dialogue content, wasting computing resources and network bandwidth, and seriously affecting the user experience. Especially when generating long content, frequent interruptions and restarts make it difficult for streaming AI dialogue systems to meet the needs of practical applications. Summary of the Invention
[0004] The main objective of this invention is to solve the technical problem that existing streaming AI dialogue systems cannot effectively resume the dialogue process after a network interruption, resulting in the loss of dialogue content. This invention provides an intelligent dialogue method based on streaming AI, the intelligent dialogue method based on streaming AI comprising: Based on the correlation between data units in the data stream output by streaming AI, a set of key data units is determined, and a storage strategy identifier is assigned to each data unit in the set of key data units. The data units corresponding to different storage strategies are stored in different levels of cache space to obtain multi-level cache data. The transmission status of the data stream is monitored. When a transmission interruption is detected, the interruption location information and the context state information of the data stream are obtained. The interruption location information, the context state information and the session identifier are associated and stored to obtain recovery information. When transmission recovery is detected, the key data unit set and historical data units are obtained from the multi-level cached data according to the recovery information, the data stream is reconstructed based on the context state information, and a request for continued data transmission is made from the streaming AI from the interruption point. The resumed data is then merged with the reconstructed data stream.
[0005] The present invention also provides an intelligent dialogue device based on streaming AI, the intelligent dialogue device based on streaming AI comprising: The key unit determination module is used to determine the set of key data units based on the correlation between data units in the data stream output by the streaming AI, and to assign a storage strategy identifier to each data unit in the set of key data units. The multi-level caching module is used to store the data units corresponding to different storage strategy identifiers into different levels of cache space to obtain multi-level cache data. The interruption monitoring module is used to monitor the transmission status of the data stream. When a transmission interruption is detected, it obtains the interruption location information and the context status information of the data stream, and associates and stores the interruption location information, the context status information and the session identifier to obtain recovery information. The resume reconstruction module is used to, when transmission recovery is detected, obtain the key data unit set and historical data units from the multi-level cache data according to the recovery information, reconstruct the data stream based on the context state information, request resume data from the streaming AI from the interruption position, and merge the resume data with the reconstructed data stream.
[0006] The present invention also provides an intelligent dialogue device based on streaming AI, comprising: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a circuit; the at least one processor invokes the instructions in the memory to cause the intelligent dialogue device based on streaming AI to perform the steps of the above-described intelligent dialogue method based on streaming AI.
[0007] The present invention also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the steps of the above-described intelligent dialogue method based on streaming AI.
[0008] The aforementioned intelligent dialogue method, apparatus, device, and storage medium based on streaming AI determine a set of key data units and assign storage strategy identifiers based on the relationships between data units in the data stream; store data units corresponding to different storage strategy identifiers in different levels of cache space; when a transmission interruption is detected, obtain the interruption location information and context state information and store them in association; when transmission recovery is detected, obtain the set of key data units and historical data units from multi-level cache data, reconstruct the data stream based on the context state information, and request resumed data from the interruption location for merging. This invention enables breakpoint resumption of streaming AI data streams, avoiding the regeneration of all content, effectively saving computing resources and network bandwidth, reducing user waiting time, and significantly improving user experience and system response efficiency.
[0009] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0010] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the first embodiment of the intelligent dialogue method based on streaming AI in this invention. Figure 2 This is a schematic diagram of a second embodiment of the intelligent dialogue method based on streaming AI in this invention. Figure 3 This is a schematic diagram of one embodiment of the intelligent dialogue device based on streaming AI in this invention. Figure 4 This is a schematic diagram of one embodiment of an intelligent dialogue device based on streaming AI in this invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0014] To facilitate understanding of this embodiment, a detailed description of an intelligent dialogue method based on streaming AI disclosed in this invention will be provided first. For example... Figure 1 As shown, this method includes the following steps: 101. Based on the correlation between data units in the data stream output by streaming AI, determine the set of key data units and assign a storage strategy identifier to each data unit in the set of key data units; In this embodiment, determining a set of key data units based on the correlation between data units in the streaming AI output data stream and assigning a storage strategy identifier to each data unit in the set of key data units includes: calculating the correlation strength of each data unit in the data stream; obtaining the cumulative correlation score of the current data unit based on the degree of dependence between the current data unit and other data units; sorting all data units in the data stream according to the cumulative correlation score and selecting the data units with the top N cumulative correlation scores as the set of key data units; and assigning a corresponding storage strategy identifier to each data unit in the set of key data units according to the score interval of the cumulative correlation score, assigning a first storage strategy identifier to data units with cumulative correlation scores in a first score interval and assigning a second storage strategy identifier to data units with cumulative correlation scores in a second score interval.
[0015] Specifically, in the process of streaming AI dialogue, data units can be tokens, phrases, or sentence fragments in the streaming output, without any specific limitations. When generating dialogue content, streaming AI will output these data units one by one in chronological order. For example, when generating a technical description, streaming AI will first output the phrase "in this invention," followed by data units such as "adopts" and "deep learning." There are semantic dependencies and logical connections between the data units.
[0016] It should be noted that the association strength calculation in this embodiment can be based on the positional relationship and semantic relevance of data units in the dialogue stream. For example, for any two data units in the data stream, the association strength value can be calculated based on their positional distance, co-occurrence frequency, and semantic relevance. Furthermore, by constructing an association strength matrix, where each row corresponds to a data unit and each column also corresponds to a data unit, the element values in the matrix represent the association strength value between the two data units.
[0017] In one embodiment, the above-mentioned calculation of the association strength for each data unit in the data stream, and the obtaining of the cumulative association score of the current data unit based on the degree of dependence between the current data unit and other data units, may include: analyzing the positional relationship and semantic association degree between any two data units in the data stream, calculating the association strength value between the two data units, and obtaining an association strength matrix; for each data unit in the data stream, extracting the association strength value between the data unit and all other data units from the association strength matrix, and summing all the extracted association strength values to obtain the cumulative association score of the data unit.
[0018] Specifically, for example, assuming the data stream contains M data units, denoted as unit1, unit2, ..., unitM, the association strength matrix can be represented as an M×M matrix, where the element in the i-th row and j-th column represents the association strength value between uniti and unitj. For a given data unit uniti, its cumulative association score can be obtained by summing all elements in the i-th row of the association strength matrix. This cumulative score reflects the overall association degree between this data unit and all other data units in the dialogue stream.
[0019] Understandably, data units located at the beginning of a data stream are usually associated with more data units because streaming AI frequently backtracks to the initial token to maintain the coherence of the generated content. For example, in a technical description, the initial subject or keywords are often repeatedly cited and associated throughout the description, so the cumulative association score of such initial data units is usually high.
[0020] After calculating the cumulative association score, all data units in the data stream can be sorted from highest to lowest score. Then, the data units with the highest cumulative association scores in the sorted results (ranked in the top N) are selected as the key data unit set. The value of N can be configured according to the actual application scenario, for example, it can be set to 4, 8, or 16. In a specific embodiment, the value of N can be dynamically adjusted according to the length and complexity of the dialogue content; a smaller value of N can be selected for shorter dialogues, and a larger value of N can be selected for longer and more complex dialogues.
[0021] After identifying the set of critical data units, storage policy identifiers need to be assigned to these data units. Specifically, multiple score intervals can be pre-defined. For example, the first score interval can be set as the interval where the cumulative score is greater than a preset threshold T1, and the second score interval can be set as the interval where the cumulative score is between T2 and T1, where T1 is greater than T2. Then, for data units whose cumulative correlation score falls within the first score interval, a first storage policy identifier is assigned. This identifier indicates that these data units have the highest importance and need to be stored in all levels of cache space. For data units whose cumulative correlation score falls within the second score interval, a second storage policy identifier is assigned. This identifier indicates that these data units have medium importance and can be selectively stored in a portion of the cache space based on a time window.
[0022] Furthermore, the step of calculating the association strength for each data unit in the data stream and obtaining the cumulative association score of the current data unit based on the dependency between the current data unit and other data units includes: analyzing the positional relationship and semantic association between any two data units in the data stream, calculating the association strength value between the two data units, and obtaining an association strength matrix; for each data unit in the data stream, extracting the association strength value between the data unit and all other data units from the association strength matrix, and summing all the extracted association strength values to obtain the cumulative association score of the data unit.
[0023] Specifically, when calculating the association strength between any two data units, it is necessary to determine both the positional relationship and the semantic association degree. The positional relationship can be determined by recording the timestamp and sequence number of each data unit in the data stream. Streaming AI, when outputting data units, marks each data unit with its positional information throughout the dialogue stream. This positional information includes the time the data unit was generated and its order in the data stream sequence. Furthermore, the positional identifiers of the two data units can be extracted, and the positional interval between them can be calculated.
[0024] In one embodiment, data units with small positional intervals typically mean that they appear sequentially during dialogue generation and have a direct contextual relationship. For example, when streaming AI generates a complete sentence, the subject, predicate, object, and other components are output sequentially. These components have small positional intervals and obvious grammatical and semantic dependencies. Different positional relevance values can be assigned to different pairs of data units based on the size of the positional interval; the smaller the positional interval, the higher the positional relevance value.
[0025] It should be noted that the calculation of positional relevance values can be done in a gradually decreasing manner. Higher positional relevance values are assigned to directly adjacent data units; as the positional interval increases, the positional relevance values gradually decrease. This decreasing rate can be adjusted according to the characteristics of the dialogue content. For technical dialogues with strong logic and close connections between sentences, a faster decreasing rate can be used; for more loosely structured everyday dialogues, a slower decreasing rate can be used.
[0026] Besides positional relationships, calculating semantic relevance is also a crucial step in determining the strength of the association. During the process of generating data units using streaming AI, each data unit is encoded as a feature representation containing semantic information. This feature representation is typically a high-dimensional vector, with each dimension corresponding to different semantic attributes, such as part-of-speech, semantic category, and contextual meaning.
[0027] In one specific embodiment, the semantic similarity of two data units can be compared by extracting their respective feature vectors. Specifically, the two feature vectors can be compared in a high-dimensional space to determine their closeness in the semantic space. If the feature vectors of two data units exhibit similar values across multiple semantic dimensions, it indicates that they are highly semantically related, such as both involving the same topic word or expressing the same type of semantic relationship. Conversely, if the feature vectors show significant differences in values across most dimensions, it indicates that the two data units are semantically weakly related.
[0028] Understandably, in actual streaming conversations, semantic associations are often not limited to lexical similarity but also include conceptual connections. For example, the data units "machine learning" and "training data," while not identical on a lexical level, are closely related conceptually because machine learning cannot function without training data. Therefore, when calculating semantic association, it is necessary to consider the conceptual connections and logical dependencies between data units.
[0029] After obtaining the positional relevance and semantic relevance values, the two can be fused. Specifically, weight coefficients can be assigned to positional relevance and semantic relevance respectively, and then a weighted combination can be performed. The weight settings can be flexibly adjusted according to the needs of the application scenario. For scenarios that emphasize temporal coherence, the weight of positional relevance can be increased; for scenarios that emphasize thematic consistency, the weight of semantic relevance can be increased.
[0030] Furthermore, by performing the above calculations on all possible data unit pairs in the data stream, a complete association strength matrix can be constructed. The construction process involves traversing each data unit in the data stream and calculating the association strength value between that data unit and all other data units. For example, for the first data unit in the data stream, it is necessary to calculate its association strength value with all other data units; for the second data unit, the same calculation is required, and so on.
[0031] After obtaining the association strength matrix, for each data unit in the data stream, all association strength values related to that data unit can be extracted from the matrix. Specifically, all values in the corresponding row of the matrix can be extracted; these values reflect the degree of association between that data unit and other data units. Then, all the extracted association strength values are summed, and the sum is the cumulative association score of that data unit. The cumulative association score reflects the overall importance of a data unit in the entire dialogue flow; data units with higher scores typically play a crucial connecting and supporting role in the dialogue.
[0032] 102. Store the data units corresponding to different storage strategy identifiers into different levels of cache space to obtain multi-level cache data; In this embodiment, storing data units corresponding to different storage policy identifiers into different levels of cache space to obtain multi-level cached data includes: obtaining the storage policy identifier and timestamp information of each data unit; determining whether the corresponding data unit is within a preset time window based on the time difference between the timestamp information and the current time; storing data units carrying the first storage policy identifier into the first-level cache space, the second-level cache space, and the third-level cache space; storing data units carrying the second storage policy identifier and located within the preset time window into the first-level cache space and the second-level cache space; and storing data units carrying the second storage policy identifier and not located within the preset time window into the third-level cache space.
[0033] Specifically, when obtaining the storage policy identifier for each data unit, the identifier information assigned to that data unit during the aforementioned key data unit identification process can be read. This identifier information indicates the importance level of the data unit, with different identifiers corresponding to different storage policies. Simultaneously, each data unit is timestamped upon generation, marking the moment the data unit was generated within the dialogue stream.
[0034] It should be noted that the time window is used to distinguish between recent and historical data units. The current time information is obtained, and then the time difference between this time and the data unit's timestamp is calculated. This time difference is then compared to a preset time window threshold. If the time difference is less than the threshold, the data unit is considered to be within the preset time window; if the time difference is greater than or equal to the threshold, the data unit is considered to be outside the time window.
[0035] In one embodiment, the size of the preset time window can be configured according to the actual situation of the dialogue. For rapidly changing dialogue scenarios, the time window can be set to be smaller; for slowly changing dialogue scenarios, the time window can be set to be larger.
[0036] Understandably, different levels of cache space differ in access speed and storage capacity. Level 1 cache space typically offers the fastest access speed but has the smallest capacity; examples include client-side local cache or browser storage. Level 2 cache space offers moderate access speed and capacity; it can be a P2P distributed cache network where nodes can share cached data. Level 3 cache space offers relatively slower access speed but larger capacity; it can be a database cache, such as a Redis cluster or persistent storage devices like MongoDB.
[0037] In one specific embodiment, the data unit carrying the first storage policy identifier needs to be stored in all levels of cache space because it has the highest importance in the conversation. Specifically, the data unit can first be written to the local cache so that it can be accessed quickly; at the same time, a copy of the data unit is distributed to several nodes of the P2P caching network; furthermore, the data unit also needs to be written to the database cache to ensure that even if the first two levels of cache fail, the data unit can still be recovered from the database.
[0038] For data units carrying a second storage policy identifier, their storage policy is differentiated based on the time window determination result. If the data unit is within the preset time window, it indicates that its generation time is relatively recent and it may be frequently accessed during dialogue resumption; therefore, it needs to be stored in the local cache and P2P cache network. If the data unit is not within the preset time window, it indicates that it has become a historical data unit and has a low probability of being accessed during dialogue resumption; therefore, it only needs to be stored in the database cache. This saves the capacity of the local cache and P2P cache, reserving space for more important or more recent data units that require faster access.
[0039] By employing the multi-level caching strategy described above, all data units in the data stream can be distributed and stored in different levels of cache space according to their importance and timeliness, resulting in multi-level cached data. This hierarchical storage method ensures both fast access and reliability of critical data units, while also making efficient use of the capacity of each level of cache space.
[0040] Furthermore, before acquiring the storage policy identifier and timestamp information of each data unit, the method further includes: monitoring network latency, network bandwidth, and packet loss rate during data stream transmission to acquire network quality parameters; calculating a network quality score based on the network quality parameters, comparing the network quality score with a preset score threshold to determine the current network quality level; adjusting the window size of the preset time window and the storage capacity allocation ratio of each level of cache space according to the quality level; when the quality level is high, decreasing the window size and increasing the allocation ratio of the third-level cache space; when the quality level is low, increasing the window size and increasing the allocation ratio of the first-level cache space.
[0041] Specifically, during streaming AI dialogue transmission, network conditions directly impact the stability and continuity of the data stream. To dynamically adjust caching strategies based on real-time network conditions, continuous monitoring of network quality is necessary. Network latency can be monitored by recording the time elapsed from data packet transmission to reception; this time interval reflects the network's response speed. Network bandwidth can be monitored by statistically analyzing the amount of data transmitted per unit time; this value reflects the network's data transmission capacity. Packet loss rate can be calculated by comparing the total number of packets sent with the number of packets successfully received; this ratio reflects the reliability of network transmission.
[0042] In one embodiment, the three network quality parameters mentioned above can be collected periodically during data stream transmission. For example, a network quality check can be performed every few seconds, recording the average network latency, average network bandwidth, and packet loss rate during that time period. These parameters can be collected by sending probe packets between the client and server. The sending and receiving times of the probe packets can be used to calculate latency, the transmission volume of the probe packets can be used to assess bandwidth, and the loss status of the probe packets can be used to calculate the loss rate.
[0043] It's important to note that calculating the network quality score requires comprehensive consideration of the three parameters mentioned above. Weighting coefficients can be assigned to network latency, network bandwidth, and packet loss rate, followed by a weighted calculation. Specifically, lower network latency, higher network bandwidth, and lower packet loss rate should result in a higher network quality score. During calculation, each parameter can first be normalized to ensure comparisons and calculations on the same scale for parameters with different dimensions. Then, the normalized parameters are weighted and summed according to their respective weighting coefficients to obtain the final network quality score.
[0044] In one specific embodiment, multiple scoring thresholds can be preset to classify network quality into different levels. For example, two dividing points can be set: a high-quality threshold and a low-quality threshold. When the network quality score is higher than the high-quality threshold, the current network is considered to be at a high-quality level; when the network quality score is lower than the low-quality threshold, the current network is considered to be at a low-quality level; and when the network quality score is between the two thresholds, the current network is considered to be at a medium-quality level.
[0045] Understandably, different network quality levels imply varying data stream transmission stability, necessitating different caching strategies. When the network is at a high quality level, data stream transmission is relatively stable with a low probability of interruption, allowing for a smaller preset time window size. This is because, under stable network conditions, it's unnecessary to retain too many recent data units in the local cache; more reliance can be placed on remote database caching. Simultaneously, the allocation ratio of the third-level cache space can be increased, storing more data units in persistent storage such as databases, thereby freeing up space in the local cache and P2P cache for other purposes.
[0046] Conversely, when the network is of low quality, data transmission is unstable and the probability of interruption is high. In this case, it is necessary to increase the size of the preset time window. This is because, under unstable network conditions, more recent data units need to be retained in the local cache to ensure rapid resumption of communication after a network outage. Simultaneously, the allocation ratio of the first-level cache space needs to be increased, storing more data units in the fastest-access local cache, reducing reliance on the remote database, and avoiding data retrieval failures due to network issues.
[0047] It's important to note that the storage capacity allocation ratio of the cache space can be adjusted through dynamic memory resource allocation. For example, when it's necessary to increase the allocation ratio of the first-level cache space, more memory space can be allocated to the local cache, while correspondingly reducing the resources allocated to other levels of cache. This dynamic adjustment method allows the caching strategy to be optimized in real time according to changes in network conditions, ensuring the stability and continuity of the conversation under different network environments.
[0048] 103. Monitor the transmission status of the data stream. When a transmission interruption is detected, obtain the interruption location information and the context status information of the data stream, and associate and store the interruption location information, the context status information and the session identifier to obtain recovery information. In this embodiment, monitoring the data stream transmission status can be achieved by detecting the connection status between the client and the streaming AI server. Specifically, heartbeat packets can be sent periodically during data stream transmission, and the connection status can be determined by the response to the heartbeat packets. If no heartbeat packet response is received within a preset time interval, or if there is an abnormal pause in data stream reception, it can be determined that the transmission is interrupted.
[0049] It's important to note that obtaining interruption location information requires recording the specific position the data stream had reached at the moment of interruption. For example, the sequence number or identifier of the last successfully received data unit can be recorded, indicating the progress of the conversation stream at the moment of interruption. Simultaneously, the character offset or token position of that data unit within the data stream can also be recorded; this information allows for precise location of the interruption.
[0050] In one embodiment, the acquisition of context state information is for the purpose of preserving the complete state of the data stream at the moment of interruption. This state information may include feature representations of key data units, correlation parameters between data units, and position encoding information. Specifically, feature vectors of each data unit can be extracted from the aforementioned stored set of key data units; these feature vectors contain the semantic information of the data units. Simultaneously, correlation weights between key data units and other data units can be extracted; these weights reflect the dependencies of data units within the dialogue stream. Furthermore, the position encoding of each data unit needs to be preserved, which marks the relative position of the data unit within the dialogue sequence.
[0051] As is understandable, a session identifier is a number or string used to uniquely identify a complete dialogue. A unique session identifier is generated for each streaming AI dialogue at the start, and this identifier remains unchanged throughout the dialogue. When a transmission interruption occurs, the session identifier of the current dialogue can be retrieved and associated with the interruption location information and context state information.
[0052] In one specific embodiment, the associated storage process can organize the above three types of information into structured data records. These data records can be in key-value pair format, using the session identifier as the index key and storing the interruption location information and context state information as the corresponding values. Specifically, a storage record for recovery information can be created in local storage space or a P2P caching network, ensuring that this record can be quickly retrieved and read during transmission recovery. This associated storage method allows for rapid location of the corresponding interruption location and context state based on the session identifier during session recovery, thereby achieving precise breakpoint resumption.
[0053] 104. When transmission recovery is detected, the key data unit set and historical data units are obtained from the multi-level cache data according to the recovery information, the data stream is reconstructed based on the context state information, and the interrupted data is requested from the streaming AI to resume transmission. The resumed data is then merged with the reconstructed data stream.
[0054] In this embodiment, transmission recovery can be detected by re-establishing the connection between the client and the streaming AI server. Specifically, when the network connection is restored or the user reopens the dialog interface, a new connection can be attempted to be established with the server. If the connection is successfully established and data can be received normally, transmission is considered to have been restored.
[0055] It should be noted that after transmission is restored, the first step is to retrieve the corresponding interruption location information and context state information from storage based on the session identifier in the recovery information. Then, data units can be obtained from multi-level cached data based on this information. Specifically, the set of key data units can be queried first from the first-level local cache. These data units are stored in the cache space at all levels in the aforementioned caching strategy, thus allowing for fast retrieval. Simultaneously, historical data units prior to the interruption location also need to be retrieved. These data units may be distributed across different levels of cache space and can be queried and retrieved sequentially from the first level to the third level.
[0056] In one embodiment, the data stream reconstruction process requires utilizing contextual state information to recover the relationships between data units. Specifically, feature vectors and association weight matrices of key data units can be extracted from the contextual state information. Then, these data units are sorted according to their positional encoding information, allowing them to be arranged in their original chronological order. Furthermore, the dependencies between data units are reconstructed based on the association weight matrix, ensuring that the reconstructed data stream maintains semantic coherence and logical integrity.
[0057] Understandably, the reconstructed data stream only contains content up to the point of interruption. To obtain the complete dialogue content, it's necessary to continue receiving new data from the point of interruption. Specifically, a resume request can be sent to the streaming AI, containing the session identifier and the interruption location information. Upon receiving the resume request, the streaming AI will resume generating dialogue content from the point of interruption and output the resumed data.
[0058] In one specific embodiment, the merging of the resumed data and the reconstructed data stream can be achieved by concatenating them in positional order. The reconstructed data stream contains all data units from the start of the dialogue to the point of interruption, while the resumed data contains newly generated data units from the point of interruption to the current moment. By concatenating the two according to their positional identifiers, a complete data stream from the start of the dialogue to the current moment can be obtained, thereby achieving seamless recovery of the dialogue content.
[0059] In this embodiment, a set of key data units is determined based on the correlation between data units in the data stream, and storage strategy identifiers are assigned. Data units corresponding to different storage strategy identifiers are stored in different levels of cache space. When a transmission interruption is detected, the interruption location information and context state information are obtained and stored in association. When transmission recovery is detected, the set of key data units and historical data units are obtained from the multi-level cache data, the data stream is reconstructed based on the context state information, and the resumed data is requested from the interruption location for merging. This invention enables breakpoint resumption of streaming AI data streams, avoids regenerating all content, effectively saves computing resources and network bandwidth, reduces user waiting time, and significantly improves user experience and system response efficiency.
[0060] Please see Figure 2 Another embodiment of the intelligent dialogue method based on streaming AI in this application includes: 201. Based on the correlation between data units in the data stream output by streaming AI, determine the set of key data units and assign a storage strategy identifier to each data unit in the set of key data units; 202. Store the data units corresponding to different storage strategy identifiers into different levels of cache space to obtain multi-level cache data; 203. Monitor the transmission status of the data stream. When a transmission interruption is detected, obtain the interruption location information and the context status information of the data stream, and associate and store the interruption location information, the context status information and the session identifier to obtain recovery information. In this embodiment, steps 201-203 are similar to steps 101-103 in the first embodiment, and will not be described again here.
[0061] 204. When transmission recovery is detected, according to the session identifier in the recovery information, data units are queried sequentially from the first-level cache space, the second-level cache space, and the third-level cache space to obtain the set of key data units and historical data units before the interruption position; In this embodiment, upon detection of transmission recovery, a session identifier can be read from the stored recovery information. This session identifier is used to retrieve the corresponding data unit in the multi-level cache space. Specifically, the query process can proceed sequentially from the first level to the third level, based on the access speed differences between different levels of cache space. The first-level cache space is typically a local cache with the fastest access speed, thus it is prioritized for querying. The second-level cache space can be a P2P caching network with moderate access speed, serving as a supplement to the first-level cache. The third-level cache space is typically a database cache with relatively slower access speed but the largest capacity, serving as the final data retrieval method.
[0062] It should be noted that during the query process, the key data unit set can be retrieved first from the first-level cache space. Since the key data units are stored in the cache spaces of all levels in the aforementioned storage strategy, they can be quickly read from the first-level cache. If the key data units with the corresponding session identifier exist in the first-level cache, these data units are directly retrieved; if they are not found in the first-level cache or the data is incomplete, the query continues to the second-level cache space.
[0063] In one embodiment, the acquisition of historical data units also follows a query order from the first level to the third level. Historical data units refer to all data units from the start of the dialogue to the point of interruption. These data units may be distributed across different levels of cache space based on their importance and timeliness. Specifically, the range of historical data units to be acquired can be determined based on the interruption location information in the recovery information, and then these data units are queried sequentially in each level of cache space. Data units found in the first-level cache are directly added to the result set; for data units not found in the first-level cache, the query continues in the second and third-level cache spaces until all required historical data units are acquired.
[0064] 205. Sort and concatenate the key data unit set and the location identifiers of the historical data units, and combine them with the context state information to obtain the reconstructed data stream; In this embodiment, the step of sorting and concatenating the key data unit set and the location identifiers of the historical data units, combined with the context state information, to obtain the reconstructed data stream includes: extracting the feature vector and location encoding information of each data unit in the key data unit set, and the association weight matrix between the key data unit set and the historical data units from the context state information; sorting the key data unit set and the historical data units in chronological order according to the location identifiers, using the key data unit set as an anchoring benchmark, and calculating the association strength value between each data unit in the historical data units and the key data unit set based on the association weight matrix; for data units with an association strength value greater than a preset threshold, retaining the association relationship between the data unit and the corresponding key data unit; for data units with an association strength value less than or equal to the preset threshold, re-establishing the association relationship between the data unit and adjacent data units based on the location encoding information to obtain the reconstructed data stream.
[0065] Specifically, the context state information stores the complete feature representation of key data units. When extracting feature vectors, the high-dimensional feature vector corresponding to each key data unit can be read from this state information. These vectors contain the semantic information and contextual features of the data unit. Simultaneously, the position encoding information marks the position of the data unit in the original dialogue stream; this encoding can be the data unit's sequence number, timestamp, or relative position offset. Furthermore, the association weight matrix records the degree of association between key data units and historical data units during dialogue generation; each element in the matrix represents the dependency strength between the two data units.
[0066] It's important to note that the sorting process rearranges the data units retrieved from the cache according to their order of appearance in the original dialogue stream. This involves reading the position identifier of each data unit, such as a timestamp or sequence number, and then sorting them in ascending order based on the size of these identifiers. This ensures that critical and historical data units are restored to their correct positions within the dialogue stream. After sorting, the entire data sequence presents a chronological order from the start of the dialogue to its interruption.
[0067] In one embodiment, a set of key data units is used as an anchoring benchmark because these data units play a core supporting role in the dialogue flow. When streaming AI generates dialogue content, it frequently revisits these key data units to maintain thematic coherence and semantic consistency. For example, in a discussion about a technical solution, core conceptual terms are often cited and associated in multiple places; these core conceptual terms correspond to key data units. By using them as anchoring benchmarks, the relationships between other data units and these key units can be reconstructed using these key units as reference points.
[0068] In one specific embodiment, the association strength value can be directly read from the association weight matrix. This matrix has been stored in the context state information before the dialogue is interrupted. Each row in the matrix corresponds to a historical data unit, and each column corresponds to a key data unit. The values of the matrix elements represent the association strength between the two. For a given historical data unit, all values in its corresponding row can be extracted from the matrix. These values represent the association strength between the data unit and each key data unit. Furthermore, the association strength value with the largest value can be selected as the association strength value of the historical data unit.
[0069] Understandably, the magnitude of the association strength value reflects the degree of dependency between historical data units and key data units. If the association strength value of a historical data unit is greater than a preset threshold, it indicates that the data unit has a strong semantic association with one or more key data units, and this association needs to be preserved when reconstructing the data stream. Specifically, the dependency relationship between the historical data unit and the corresponding key data unit can be marked in the reconstructed data stream, so that the semantic information of the relevant key data units can be referenced when understanding the historical data unit.
[0070] For data units with an association strength value less than or equal to a preset threshold, it indicates that their direct association with the key data unit is weak, and they rely more on their adjacent data units. In this case, the association can be re-established based on location encoding information. Specifically, the adjacent data units before and after the data unit can be found based on the location encoding, and then the dependency relationship between the data unit and its adjacent data units can be established. This location-based association establishment method can ensure the coherence of the data flow within a local scope. Even if some data units have a weak global association with the key data unit, semantic continuity can still be maintained through local adjacency relationships.
[0071] 206. Request resumed data from the interrupted point in the streaming AI, and merge the resumed data with the reconstructed data stream.
[0072] In this embodiment, a set of key data units is determined based on the correlation between data units in the data stream, and storage strategy identifiers are assigned. Data units corresponding to different storage strategy identifiers are stored in different levels of cache space. When a transmission interruption is detected, the interruption location information and context state information are obtained and stored in association. When transmission recovery is detected, the set of key data units and historical data units are obtained from the multi-level cache data, the data stream is reconstructed based on the context state information, and the resumed data is requested from the interruption location for merging. This invention enables breakpoint resumption of streaming AI data streams, avoids regenerating all content, effectively saves computing resources and network bandwidth, reduces user waiting time, and significantly improves user experience and system response efficiency.
[0073] The above describes the intelligent dialogue method based on streaming AI in the embodiments of the present invention. The following describes the intelligent dialogue device based on streaming AI in the embodiments of the present invention. Please refer to [link to relevant documentation] for details on this intelligent dialogue device. Figure 3 One embodiment of the intelligent dialogue device based on streaming AI in this invention includes: The key unit determination module 301 is used to determine the key data unit set based on the correlation between each data unit in the data stream output by the streaming AI, and to assign a storage strategy identifier to each data unit in the key data unit set. The multi-level cache module 302 is used to store the data units corresponding to different storage strategy identifiers into different levels of cache space to obtain multi-level cache data. The interruption monitoring module 303 is used to monitor the transmission status of the data stream. When a transmission interruption is detected, it acquires the interruption location information and the context status information of the data stream, and associates and stores the interruption location information, the context status information and the session identifier to obtain recovery information. The resume reconstruction module 304 is used to, when a transmission recovery is detected, obtain the key data unit set and historical data units from the multi-level cache data according to the recovery information, reconstruct the data stream based on the context state information, request resume data from the streaming AI from the interruption position, and merge the resume data with the reconstructed data stream.
[0074] In this embodiment of the invention, the intelligent dialogue device based on streaming AI runs the aforementioned intelligent dialogue method based on streaming AI. The device determines a set of key data units and assigns storage strategy identifiers based on the relationships between data units in the data stream; it stores data units corresponding to different storage strategy identifiers into different levels of cache space; when a transmission interruption is detected, it obtains the interruption location information and context state information and stores them in association; when transmission recovery is detected, it obtains the set of key data units and historical data units from the multi-level cache data, reconstructs the data stream based on the context state information, and requests resumed data from the interruption location for merging. This invention enables breakpoint resumption of streaming AI data streams, avoiding the regeneration of all content, effectively saving computing resources and network bandwidth, reducing user waiting time, and significantly improving user experience and system response efficiency.
[0075] above Figure 3 The intelligent dialogue device based on streaming AI in the embodiments of the present invention will be described in detail from the perspective of unitized functional entities. The intelligent dialogue device based on streaming AI in the embodiments of the present invention will be described in detail from the perspective of hardware processing.
[0076] Figure 4This is a schematic diagram of the structure of a streaming AI-based intelligent dialogue device 400 provided in an embodiment of the present invention. The streaming AI-based intelligent dialogue device 400 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 410 (e.g., one or more processors) and a memory 420, and one or more storage media 430 (e.g., one or more mass storage devices) for storing application programs 433 or data 432. The memory 420 and storage media 430 can be temporary or persistent storage. The program stored in the storage media 430 may include one or more units (not shown in the diagram), each unit may include a series of instruction operations on the streaming AI-based intelligent dialogue device 400. Furthermore, the processor 410 may be configured to communicate with the storage media 430 and execute a series of instruction operations in the storage media 430 on the streaming AI-based intelligent dialogue device 400 to implement the steps of the above-described streaming AI-based intelligent dialogue method.
[0077] The streaming AI-based intelligent dialogue device 400 may also include one or more power supplies 440, one or more wired or wireless network interfaces 450, one or more input / output interfaces 460, and / or one or more operating systems 431, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 4 The illustrated structure of the AI-based intelligent dialogue device does not constitute a limitation on the AI-based intelligent dialogue device provided by this invention. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0078] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the intelligent dialogue method based on streaming AI.
[0079] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0080] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0081] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A smart dialogue method based on streaming AI, characterized in that, The intelligent dialogue method based on streaming AI includes: Based on the correlation between data units in the data stream output by streaming AI, a set of key data units is determined, and a storage strategy identifier is assigned to each data unit in the set of key data units. The data units corresponding to different storage strategies are stored in different levels of cache space to obtain multi-level cache data. The transmission status of the data stream is monitored. When a transmission interruption is detected, the interruption location information and the context state information of the data stream are obtained. The interruption location information, the context state information and the session identifier are associated and stored to obtain recovery information. When transmission recovery is detected, the key data unit set and historical data units are obtained from the multi-level cached data according to the recovery information, the data stream is reconstructed based on the context state information, and a request for continued data transmission is made from the streaming AI from the interruption point. The resumed data is then merged with the reconstructed data stream.
2. The intelligent dialogue method based on streaming AI according to claim 1, characterized in that, The step of determining a set of key data units based on the correlation between data units in the streaming AI output data stream, and assigning a storage strategy identifier to each data unit in the set of key data units, includes: For each data unit in the data stream, the association strength is calculated, and the cumulative association score of the current data unit is obtained based on the degree of dependence between the current data unit and other data units. All data units in the data stream are sorted according to the cumulative association score, and the data units with the top N cumulative association scores are selected as the key data unit set. Based on the score interval where the cumulative correlation score is located, each data unit in the key data unit set is assigned a corresponding storage strategy identifier. Data units whose cumulative correlation scores are located in the first score interval are assigned a first storage strategy identifier, and data units whose cumulative correlation scores are located in the second score interval are assigned a second storage strategy identifier.
3. The intelligent dialogue method based on streaming AI according to claim 2, characterized in that, The step of calculating the association strength for each data unit in the data stream, and obtaining the cumulative association score of the current data unit based on the dependency between the current data unit and other data units, includes: The positional relationship and semantic correlation between any two data units in the data stream are analyzed, and the correlation strength value between the two data units is calculated to obtain the correlation strength matrix. For each data unit in the data stream, the association strength value between the data unit and all other data units is extracted from the association strength matrix. The extracted association strength values are summed to obtain the cumulative association score of the data unit.
4. The intelligent dialogue method based on streaming AI according to claim 2, characterized in that, The step of storing the data units corresponding to different storage strategy identifiers into different levels of cache space to obtain multi-level cache data includes: Obtain the storage strategy identifier and timestamp information for each data unit, and determine whether the corresponding data unit is within a preset time window based on the time difference between the timestamp information and the current time. The data unit carrying the first storage strategy identifier is stored in the first-level cache space, the second-level cache space, and the third-level cache space; Data units carrying a second storage strategy identifier and located within the preset time window are stored in the first-level cache space and the second-level cache space; Data units carrying the second storage strategy identifier and not located within the preset time window are stored in the third-level cache space.
5. The intelligent dialogue method based on streaming AI according to claim 4, characterized in that, Before obtaining the storage policy identifier and timestamp information for each data unit, the method further includes: Monitor network latency, network bandwidth, and packet loss rate during data stream transmission to obtain network quality parameters; The network quality score is calculated based on the network quality parameters, and the network quality score is compared with a preset score threshold to determine the current quality level of the network. The window size of the preset time window and the storage capacity allocation ratio of each level of cache space are adjusted according to the quality level. When the quality level is high, the window size is reduced and the allocation ratio of the third-level cache space is increased. When the quality level is low, the window size is increased and the allocation ratio of the first-level cache space is increased.
6. The intelligent dialogue method based on streaming AI according to claim 1, characterized in that, The step of obtaining the key data unit set and historical data units from the multi-level cache data based on the recovery information, and reconstructing the data stream based on the context state information includes: Based on the session identifier in the recovery information, data units are queried sequentially from the first-level cache space, the second-level cache space, and the third-level cache space to obtain the set of key data units and historical data units before the interruption position; The data is sorted and concatenated based on the location identifiers of the key data units and the historical data units, and combined with the context state information to obtain the reconstructed data stream.
7. The intelligent dialogue method based on streaming AI according to claim 5, characterized in that, The step of sorting and concatenating the key data unit set and the location identifiers of the historical data units, combined with the context state information, to obtain the reconstructed data stream includes: Extract the feature vector and position encoding information of each data unit in the key data unit set from the context state information, as well as the association weight matrix between the key data unit set and historical data units; The key data unit set and the historical data unit are sorted in chronological order according to the location identifier. The key data unit set is used as the anchoring benchmark. The association strength value between each data unit in the historical data unit and the key data unit set is calculated according to the association weight matrix. For data units with an association strength value greater than a preset threshold, the association relationship between the data unit and the corresponding key data unit is retained; For data units whose association strength value is less than or equal to the preset threshold, the association relationship between the data unit and its neighboring data units is re-established based on the location encoding information to obtain a reconstructed data stream.
8. A smart dialogue device based on streaming AI, characterized in that, The intelligent dialogue device based on streaming AI includes: The key unit determination module is used to determine the set of key data units based on the correlation between data units in the data stream output by the streaming AI, and to assign a storage strategy identifier to each data unit in the set of key data units. The multi-level caching module is used to store the data units corresponding to different storage strategy identifiers into different levels of cache space to obtain multi-level cache data. The interruption monitoring module is used to monitor the transmission status of the data stream. When a transmission interruption is detected, it obtains the interruption location information and the context status information of the data stream, and associates and stores the interruption location information, the context status information and the session identifier to obtain recovery information. The resume reconstruction module is used to, when transmission recovery is detected, obtain the key data unit set and historical data units from the multi-level cache data according to the recovery information, reconstruct the data stream based on the context state information, request resume data from the streaming AI from the interruption position, and merge the resume data with the reconstructed data stream.
9. A smart dialogue device based on streaming AI, characterized in that, The intelligent dialogue device based on streaming AI includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the streaming AI-based intelligent dialogue device to perform the steps of the streaming AI-based intelligent dialogue method as described in any one of claims 1-7.
10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the steps of the intelligent dialogue method based on streaming AI as described in any one of claims 1-7.