A resource coordination scheduling method and device of intelligent glasses, medium and product

CN122816795APending Publication Date: 2026-09-25SHENZHEN XEME COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610951523.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

纯端侧方案的缺点在于,受限于端侧芯片的算力和存储资源,压缩后的小模型的性能有限,难以有效处理复杂的逻辑推理、长上下文的连续对话或多模态融合任务,导致用户体验下降

Benefits of technology

1、智能眼镜通过资源协同调度方法,实现了对语音指令处理的动态适配。通过综合分析指令复杂度与上下文复杂度,并基于最终执行复杂度选择预设处理模式(包括纯本地极速模式、本地切片混合模式和端云协同模式),有效平衡了系统实时性、处理性能与功耗之间的矛盾,显著提升了用户体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816795A_ABST
    Figure CN122816795A_ABST
Patent Text Reader

Abstract

A resource cooperative scheduling method and device of smart glasses, medium and product, relate to the field of smart glasses. In the method, a voice instruction input by a user in a current dialogue turn is acquired; a linguistic feature of the voice instruction is analyzed to generate an instruction complexity, the linguistic feature including an instruction duration, keywords and a syntax structure; a historical dialogue is read and a context complexity score is determined based on the historical dialogue; a final execution complexity is determined by synthesizing the context complexity score and the instruction complexity; based on the final execution complexity, a target processing mode matching the final execution complexity is determined in a preset processing mode; a computing resource corresponding to the target processing mode is called to process the voice instruction, and a reply result is output. The technical solution provided in the application improves the utilization efficiency of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart glasses, specifically to a resource collaborative scheduling method, device, medium, and product for smart glasses. Background Technology

[0002] With the rapid development of artificial intelligence and wearable device technologies, smart glasses, as a new generation of consumer electronics products, are gradually becoming an important terminal for users to conduct real-time interaction, obtain information, and receive intelligent assistance in their daily lives, thanks to their unique advantages such as freeing up their hands and providing a first-person perspective.

[0003] Currently, the mainstream AI smart glasses technologies on the market mainly fall into two categories. The first is a pure cloud-based solution, where the smart glasses only act as a medium for audio acquisition and transmission; all complex computational tasks, especially inference using large language models, are completed on cloud servers. The drawback of this pure cloud-based solution is its extremely high dependence on network connectivity, which introduces significant network latency, affecting the real-time nature of user interaction. Furthermore, if the network signal is unstable or disconnected, the core functions of the smart glasses will completely fail. The second is a pure edge-side solution, which attempts to deploy compressed, miniaturized models directly on the edge chip of the smart glasses. The drawback of this pure edge-side solution is that, limited by the computing power and storage resources of the edge chip, the performance of the compressed mini-model is limited, making it difficult to effectively handle complex logical reasoning, long-context continuous dialogue, or multimodal fusion tasks, resulting in a degraded user experience. Existing technologies, whether pure cloud-based or pure edge-side, cannot simultaneously meet the requirements of low latency, low power consumption, and high performance, leading to low efficiency in the utilization of computing resources. Summary of the Invention

[0004] This application provides a resource collaborative scheduling method, device, medium, and product for smart glasses, which improves the efficiency of computing resource utilization.

[0005] A first aspect of this application provides a resource collaborative scheduling method for smart glasses. The method includes: acquiring a voice command input by a user in the current dialogue round; analyzing the linguistic features of the voice command to generate a command complexity, the linguistic features including command duration, keywords, and syntactic structure; reading historical dialogues and determining a context complexity score based on the historical dialogues; combining the context complexity score and the command complexity to determine a final execution complexity; based on the final execution complexity, determining a target processing mode matching the final execution complexity in a preset processing mode; and calling the computing resources corresponding to the target processing mode to process the voice command and output a response result.

[0006] By employing the aforementioned technical solutions, the system acquires the voice commands input by the user in the current dialogue round, achieving real-time capture of the user's interactive intent and ensuring that the smart glasses can respond to user needs promptly. The linguistic features of the voice commands are analyzed to generate command complexity. By comprehensively considering three dimensions—command duration, keywords, and syntactic structure—a quantitative assessment of command processing difficulty is achieved, providing a precise basis for subsequent resource scheduling decisions. Reading historical dialogues and determining contextual complexity scores based on them enables the system to understand the continuity and relevance of dialogues, avoiding semantic comprehension biases caused by processing isolated single commands. The final execution complexity is determined by combining the contextual complexity score and command complexity, achieving a multi-dimensional complexity fusion assessment and providing a more comprehensive and accurate judgment of task difficulty. Based on the final execution complexity, a target processing mode is determined from preset processing modes, achieving precise matching between processing modes and task complexity, avoiding the problem of simple tasks consuming excessive computing resources or complex tasks failing due to insufficient resources. The system calls upon the computing resources corresponding to the target processing mode to process voice commands and output response results, thereby achieving dynamic optimization of computing resource allocation. This maximizes resource utilization efficiency while ensuring processing quality and extending the battery life of smart glasses.

[0007] Optionally, the step of analyzing the linguistic features of the voice command to generate command complexity specifically includes: transcribing the voice command to obtain command text, extracting the command duration corresponding to the command text, and determining whether the command duration is less than a preset duration threshold; performing keyword matching on the command text to extract semantic features representing whether there are words that require external knowledge reasoning; performing syntactic parsing on the command text to extract syntactic structure features representing sentence nesting depth and the number of components; and fusing the command duration, the semantic features, and the syntactic structure features according to preset weights to generate the command complexity.

[0008] By employing the above technical solutions, speech commands are transcribed to obtain command text, and command duration is extracted. By determining whether the command duration is less than a preset duration threshold, a rapid initial screening of command length complexity is achieved. Short commands typically correspond to simple tasks, while long commands often contain more information requiring processing. Keyword matching is performed on the command text to extract semantic features indicating the presence of words requiring external knowledge reasoning. This enables the system to identify complex queries that require knowledge base access or deep reasoning, providing a semantic basis for resource scheduling. Syntactic parsing of the command text extracts syntactic structural features representing sentence nesting depth and the number of components, achieving precise quantification of language structural complexity. Sentences with deeper nesting and more components require more computational resources for understanding and processing. Command duration, semantic features, and syntactic structural features are fused according to preset weights to generate command complexity. Multi-feature weighted fusion achieves a comprehensive evaluation of command complexity. The preset weights allow for optimization and adjustment of the contribution of different features to the final complexity based on actual application scenarios, improving the accuracy and adaptability of complexity evaluation. This multi-dimensional feature fusion method avoids the one-sidedness of judging by a single feature and provides a more reliable quantitative basis for subsequent resource scheduling decisions.

[0009] Optionally, the step of reading historical dialogues and determining a context complexity score based on the historical dialogues specifically includes: detecting whether there are referential words or ellipsis components in the instruction text; when the referential words or ellipsis components are detected, reading the historical dialogues stored in the local cache and vectorizing the historical dialogues to obtain historical semantic vectors; calculating the semantic correlation degree between the current semantic vector corresponding to the voice instruction and the historical semantic vectors, and counting the cumulative number of rounds of the historical dialogues; and determining the context complexity score based on the semantic correlation degree and the cumulative number of rounds.

[0010] By employing the aforementioned technical solution, the system detects the presence of referential words or ellipsis in the instruction text, achieving accurate identification of contextual dependencies and determining that the current instruction requires the integration of historical dialogue for complete understanding. Upon detecting referential words or ellipsis, the system reads the historical dialogue stored in the local cache and vectorizes it to obtain a historical semantic vector. This vectorization enables numerical encoding of the semantic information of the historical dialogue, facilitating subsequent similarity calculations and association analysis. The semantic correlation between the current semantic vector and the historical semantic vector corresponding to the voice instruction is calculated. Similarity calculation in the vector space quantifies the degree of semantic correlation between the current and historical dialogues; a higher correlation indicates a greater reliance on historical context for understanding the current instruction. The cumulative number of turns in the historical dialogue is counted, and a context complexity score is determined based on the semantic correlation and the cumulative number of turns. This achieves a multi-factor comprehensive evaluation of context complexity; a higher cumulative number of turns indicates a longer dialogue chain, requiring more contextual information to be maintained, and thus higher processing complexity. This method ensures that the system can accurately assess the processing needs in multi-turn dialogue scenarios, providing an important basis for resource scheduling in continuous dialogue scenarios.

[0011] Optionally, the preset processing modes include a pure local high-speed mode, a local slicing hybrid mode, and an edge-cloud collaborative mode. The step of determining a target processing mode that matches the final execution complexity within the preset processing modes, based on the final execution complexity, specifically includes: if the final execution complexity is less than or equal to a first preset threshold, then the target processing mode is determined to be a pure local high-speed mode; if the final execution complexity is greater than the first preset threshold but less than a second preset threshold, then the target processing mode is determined to be a local slicing hybrid mode, where the second preset threshold is greater than the first preset threshold; if the final execution complexity is greater than or equal to the second preset threshold, then the target processing mode is determined to be an edge-cloud collaborative mode.

[0012] By adopting the above technical solution, three preset processing modes—pure local high-speed mode, local slicing hybrid mode, and edge-cloud collaborative mode—were set up, constructing a hierarchical processing architecture covering different complexity levels and realizing gradient configuration of computing resources. When the final execution complexity is less than or equal to the first preset threshold, the target processing mode is determined to be pure local high-speed mode, enabling simple tasks to be processed quickly locally, avoiding unnecessary model loading and network transmission overhead, achieving millisecond-level response speeds, and significantly improving user experience. When the final execution complexity is greater than the first preset threshold but less than the second preset threshold, the target processing mode is determined to be local slicing hybrid mode, achieving optimized processing of medium-complexity tasks. Partial activation of the local model avoids the resource consumption of full model loading while ensuring processing quality. When the final execution complexity is greater than or equal to the second preset threshold, the target processing mode is determined to be edge-cloud collaborative mode, ensuring that complex tasks receive sufficient computing resource support, and guaranteeing the accuracy of processing results through the powerful computing power of the cloud. The setting of the first and second preset thresholds achieves a reasonable division of the complexity space. The configurability of the thresholds allows the system to be flexibly adjusted according to different application scenarios and hardware conditions, improving the versatility and adaptability of the method.

[0013] Optionally, the step of calling the computing resources corresponding to the target processing mode to process the voice command and output a response result specifically includes: in the pure local high-speed mode, inputting the voice command to a dedicated neural processing unit on the smart glasses side for processing, and having the dedicated neural processing unit generate the response result, wherein the dedicated neural processing unit is a lightweight model; in the local slicing hybrid mode, calling the locally deployed full-scale language model and activating the model slicing engine; the model slicing engine, according to the voice command, activates and powers the first neuron layer group corresponding to simple reasoning in the full-scale language model, and cuts off the second neuron layer group related to deep reasoning. Powered by the meta-layer group, the first neuron layer group includes a understanding layer and a simple logic layer, and the second neuron layer group includes a reasoning layer and a generation layer. The first neuron layer group processes the voice command to generate the response result. In the edge-cloud collaborative mode, the intent extraction capability of the locally deployed full-scale language model is invoked to compress the voice command into compressed prompt words containing core semantic information. The compressed prompt words are uploaded to the cloud server for processing, and while waiting for the cloud to return the result, the local main computing unit of the smart glasses is placed in a deep standby state. The returned data from the cloud server is received and decoded to form the response result.

[0014] By adopting the above technical solutions, in the pure local high-speed mode, voice commands are input to a dedicated neural processing unit on the smart glasses side for processing. A lightweight model enables rapid local processing of simple tasks, and the use of a dedicated neural processing unit ensures maximized computational efficiency. The deployment of the lightweight model reduces memory usage and power consumption. In the local slicing hybrid mode, the fully deployed local large language model is invoked, and the model slicing engine is activated, enabling on-demand partial activation of the large model and avoiding the huge resource consumption of running the entire model. The model slicing engine activates the first neuron layer group and cuts off the power supply to the second neuron layer group according to the voice command. Precise hierarchical control enables fine-grained management of computing resources. The first neuron layer group, including the understanding layer and simple logic layer, ensures basic reasoning capabilities, while the power cut-off of the second neuron layer group significantly reduces power consumption. In the edge-cloud collaborative mode, the intent extraction capability of the local model compresses voice commands into compressed prompt words, significantly reducing the amount of data transmitted and lowering network bandwidth requirements and transmission latency. The compressed prompt words are uploaded to the cloud server for processing, and the local main computing unit is placed in deep standby mode during the waiting period, achieving intelligent energy saving of local resources. The deep standby state minimizes power consumption during the waiting period. Receiving and decoding data returned by the cloud server to form a response ensures that complex tasks can obtain high-quality processing results.

[0015] Optionally, the step of the model slicing engine activating and powering the first neuron layer group corresponding to simple reasoning in the full-scale language model and cutting off the power supply to the second neuron layer group related to deep reasoning, according to the voice command, specifically includes: pre-storing the complete parameters of the full-scale language model in the non-volatile memory of the smart glasses, and establishing a hierarchical index table of model parameters, the hierarchical index table including the starting storage address, parameter size, and inter-layer dependencies of each transformer layer; parsing the task type of the voice command, and querying the preset task hierarchical mapping table to determine the first neuron layer group and the second neuron layer group. A slice activation mask is generated; based on the slice activation mask and the hierarchical index table, the weight parameters corresponding to the first neuron layer group are read from the non-volatile memory through the memory access controller, and the weight parameters are transferred to the buffer of the on-chip static random access memory; a power control signal is sent to the hardware-level power management unit to apply the working voltage to the multiply-accumulate operation array corresponding to the first neuron layer group and turn on the clock signal; the clock tree of the computing unit corresponding to the second neuron layer group is turned off through clock gating technology, and the power supply rail of the computing unit corresponding to the second neuron layer group is cut off through voltage domain isolation technology, so that the second neuron layer group is in a zero static power consumption state.

[0016] By adopting the above technical solution, the complete parameters of the full large language model are pre-stored in the non-volatile memory of the smart glasses, and a hierarchical index table is established. This achieves structured storage and rapid location of model parameters. The hierarchical index table, containing information on the starting storage address, parameter size, and inter-layer dependencies, provides precise addressing capabilities for subsequent selective loading. The task type of the voice command is parsed, and the neuron layer group is determined by querying the preset task hierarchical mapping table, achieving a precise mapping between task requirements and model structure. The generated slice activation mask provides precise control information for hierarchical activation. Based on the slice activation mask and the hierarchical index table, the weight parameters corresponding to the first neuron layer group are read from the non-volatile memory through the memory access controller and transferred to the on-chip static random access memory (SRAM), achieving selective loading of parameters. The use of SRAM significantly improves parameter access speed and reduces memory access latency. A power control signal is sent to the hardware-level power management unit to apply the operating voltage to the first neuron layer group and enable the clock signal, achieving precise power supply control of the computing unit, providing power only to the parts that need to work. By shutting down the clock tree of the second neuron layer using clock gating technology and cutting off the power supply rails using voltage domain isolation technology, the unused parts are completely powered off. The zero static power consumption state maximizes the reduction of overall power consumption and significantly extends the battery life of the smart glasses.

[0017] Optionally, after calling the computing resources corresponding to the target processing mode to process the voice command and output the response result, the method further includes: determining whether there is a semantic relationship between the next round of dialogue and the current dialogue; if it is determined that there is a semantic relationship between the next round of dialogue and the current dialogue, then retaining the activated slice of the current dialogue in the cache memory; determining the difference between the target slice required for the next round of dialogue and the activated slice, and loading the difference increment into the on-chip static random access memory; maintaining the reuse statistics of each slice in the cache memory, and, according to a preset minimum usage strategy, removing slices that have not been reused in a preset number of rounds of dialogue from the cache memory.

[0018] By employing the above technical solutions, the system determines whether there is a semantic relationship between the next round of dialogue and the current dialogue, achieving intelligent recognition of dialogue continuity and providing a decision-making basis for caching strategy formulation. If a semantic relationship is determined between the next round of dialogue and the current dialogue, the activated slices of the current dialogue are retained in the cache memory, realizing intelligent reuse of model slices and avoiding the time and energy overhead caused by repeatedly loading the same slices, significantly improving the response speed in multi-round dialogue scenarios. The system determines the difference between the target slice required for the next round of dialogue and the activated slices, and incrementally loads the difference set into on-chip static random access memory. The incremental loading strategy minimizes memory movement operations, loading only the newly required slice portion, reducing memory bandwidth consumption and loading latency. Maintaining reuse statistics of each slice in the cache memory enables dynamic tracking of slice usage frequency, providing a quantitative basis for cache replacement decisions. Based on a preset minimum usage strategy, unused slices from a preset number of rounds of dialogue are removed from the cache memory, achieving dynamic optimization management of cache space, ensuring that frequently used slices can reside in the cache, and that infrequently used slices are promptly cleared and space released, maximizing the cache hit rate. This intelligent cache management mechanism effectively controls memory usage while ensuring processing performance, achieving an optimal balance between storage resources and computing efficiency.

[0019] Secondly, embodiments of this application provide a resource coordination scheduling device for smart glasses. The resource coordination scheduling device for smart glasses includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the resource coordination scheduling device for smart glasses to perform the method described in the first aspect and any possible implementation thereof.

[0020] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a resource coordination scheduling device of smart glasses, cause the resource coordination scheduling device of smart glasses to perform the method described in the first aspect and any possible implementation thereof.

[0021] Fourthly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a resource coordination scheduling device of smart glasses, causes the resource coordination scheduling device of the smart glasses to execute the method described in the first aspect and any possible implementation thereof.

[0022] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: 1. Smart glasses achieve dynamic adaptation of voice command processing through resource collaborative scheduling. By comprehensively analyzing the complexity of the command and the context, and selecting preset processing modes (including pure local high-speed mode, local slice hybrid mode, and edge-cloud collaborative mode) based on the final execution complexity, the contradiction between system real-time performance, processing performance, and power consumption is effectively balanced, significantly improving the user experience.

[0023] 2. By employing linguistic feature analysis of voice commands (such as command duration, keywords, and syntactic structure) and contextual complexity assessment (including semantic relevance and number of dialogue turns), the system can accurately determine task complexity and achieve resource matching between complex and simple tasks. Through dynamic task allocation, the system further optimizes the utilization efficiency of computing resources, enabling efficient localized processing of low-complexity tasks and cloud-based collaborative execution of high-complexity tasks.

[0024] 3. By combining model slicing technology with hardware-level power management, the local slicing hybrid mode can activate only the necessary neuron layers while cutting off power to the rest without affecting inference performance, thus reducing hardware power consumption and maintaining computational performance. Slice caching and incremental loading mechanisms support rapid response to multi-turn dialogue-related tasks, further improving the system's real-time performance and resource reuse efficiency. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating a resource collaborative scheduling method for smart glasses disclosed in an embodiment of this application; Figure 2 This is another flowchart illustrating a resource collaborative scheduling method for smart glasses disclosed in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a resource collaborative scheduling device for smart glasses provided in an embodiment of this application.

[0026] Explanation of reference numerals in the attached drawings: 301, Central Processing Unit; 302, Read-Only Memory; 303, Random Access Memory; 304, Bus; 305, Input / Output Interface; 306, Input Section; 307, Output Section; 308, Storage Section; 309, Communication Section; 310, Driver; 311, Removable Media. Detailed Implementation

[0027] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0028] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0029] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple system devices refer to two or more system devices, and multiple screen terminals refer to two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0030] This application provides a resource collaborative scheduling method for smart glasses, referring to... Figure 1 , Figure 1 This is a flowchart illustrating a resource collaborative scheduling method for smart glasses according to an embodiment of this application. The method is applied to smart glasses, enabling the smart glasses to execute a resource collaborative scheduling program. The method includes steps S101 to S105, as follows: Step S101: Obtain the voice command input by the user in the current dialogue round.

[0031] In step S101, "user" refers to the operator who uses the smart glasses and issues a voice interaction request. The current dialogue turn refers to a single interaction process from when the user begins speaking to when the smart glasses provide a complete response. A voice command refers to the raw acoustic signal issued by the user through spoken language and collected by the audio sensors of the smart glasses.

[0032] Specifically, the smart glasses monitor ambient sounds in real time via a built-in microphone array. When the system detects that the user has uttered a preset wake word or manually triggered an interaction, the smart glasses' audio capture system begins to record the user's voice command in its entirety. This voice command is then converted into a digital audio stream, serving as the raw input data for subsequent processing.

[0033] Step S102: Analyze the linguistic features of the voice command and generate the command complexity. The linguistic features include command duration, keywords, and syntactic structure.

[0034] In step S102, linguistic features refer to linguistic attributes extracted from speech or text that reflect the inherent complexity of the instruction. Instruction complexity is a numerical value obtained by quantifying linguistic features, used to measure the amount of cognitive and computational resources required to process the speech instruction. Instruction duration refers to the duration of the speech instruction from start to finish. Keywords refer to specific words contained in the instruction text that indicate the need for complex reasoning or the invocation of external knowledge. Syntactic structure refers to the sentence organization of the instruction text, such as the depth of sentence nesting and the number of grammatical components.

[0035] Specifically, the smart glasses first process the acquired voice command, measuring the duration of the audio signal to obtain the command duration. Simultaneously, the smart glasses use a local speech-to-text engine to convert the voice command into text. Next, the smart glasses match the text against a pre-defined keyword library containing words suggesting complex tasks, such as "analyze," "compare," "plan," and "why." Keyword features are determined based on the number and weight of matched keywords. Subsequently, the smart glasses perform syntactic analysis on the command text, generating a parse tree. The syntactic structure features are quantified by calculating the depth and number of nodes in this parse tree. Finally, the smart glasses weight and fuse the command duration, keyword features, and syntactic structure features according to pre-defined weights to calculate a comprehensive value, which is the command complexity.

[0036] In one possible implementation, the linguistic features of the voice command are analyzed to generate the command complexity, specifically including steps S1021-S1024, as follows: Step S1021: Transcribe the voice command to obtain the command text, extract the command duration corresponding to the command text, and determine whether the command duration is less than the preset duration threshold.

[0037] In step S1021, the instruction text refers to the text sequence obtained by converting the user's original voice signal through speech recognition technology. The preset duration threshold refers to a pre-set time length standard in the smart glasses system, used to quickly and initially assess the length of the voice instruction.

[0038] Specifically, after receiving a user's voice command, the smart glasses invoke the on-device automatic speech recognition engine to transcribe the acquired digital audio stream into text in real time. Simultaneously, the smart glasses' timer precisely measures the complete duration from the start to the end of the speech; this duration is the command duration. The smart glasses then compare the acquired command duration with a preset duration threshold, such as 2 seconds, stored locally. The comparison result—whether the command duration is less than the preset threshold—is recorded as a Boolean value for subsequent command complexity calculations.

[0039] Step S1022: Perform keyword matching on the instruction text and extract semantic features that indicate whether there are words that require external knowledge reasoning.

[0040] In step S1022, keyword matching refers to the process of searching for and identifying specific words contained in a preset thesaurus within the instruction text. External knowledge reasoning refers to the process where processing instructions requires reliance on a general knowledge base, a specialized domain database, or multi-step logical deduction. Semantic features refer to the feature information extracted through keyword matching that characterizes whether the instruction intent involves a complex task.

[0041] Specifically, the smart glasses access a locally stored keyword lexicon. This lexicon predefines a large number of words related to complex tasks, such as "analysis," "comparison," "summary," "planning," "reasons," and "evaluation." The smart glasses take the instruction text generated in step S1021 as input and search and match it in the keyword lexicon. The smart glasses count the number of matching keywords in the instruction text and assign different weights to different keywords based on their complexity. For example, "analysis" has a higher weight than "query." Finally, the smart glasses calculate a numerical value based on the number of matched keywords and their weights; this numerical value is the semantic feature representing the semantic complexity of the instruction.

[0042] Step S1023: Perform syntactic parsing on the instruction text to extract syntactic structural features that represent the nesting depth and number of components of the sentence.

[0043] In step S1023, syntactic parsing refers to analyzing the sentence structure of the instruction text to identify grammatical components such as subject, predicate, object, attributive, and adverbial, as well as the hierarchy and dependencies between these components. Sentence nesting depth refers to the longest path length from the root node to the farthest leaf node in the syntactic parsing tree, reflecting the complexity of clauses or modifying elements. The number of components refers to the total number of independent grammatical components identified after syntactic analysis. Syntactic structural features are characteristics obtained by quantifying indicators such as sentence nesting depth and the number of components, used to measure the complexity of sentence structure.

[0044] Specifically, the smart glasses utilize built-in natural language processing tools to perform syntactic parsing on the instruction text, generating a visual syntactic analysis tree. The smart glasses' algorithm traverses this syntactic analysis tree, calculating its maximum depth, which is the sentence nesting depth. Simultaneously, the algorithm counts the total number of nodes in the tree, which is the number of constituents. The smart glasses combine or weight these two values—sentence nesting depth and number of constituents—to arrive at a comprehensive value, which is the syntactic structure feature used to quantify the syntactic complexity of the instruction text.

[0045] Step S1024: Instruction duration, semantic features and syntactic structure features are fused according to preset weights to generate instruction complexity.

[0046] In step S1024, the preset weights refer to the numerical coefficients pre-assigned to the three different dimensions of features—instruction duration, semantic features, and syntactic structure features—to reflect the importance of different features in assessing the overall complexity during the final calculation. Fusion refers to the process of combining feature information from multiple different sources into a single comprehensive index through mathematical operations.

[0047] Specifically, the smart glasses obtain the result of the instruction duration judgment, the numerical values ​​of semantic features, and the numerical values ​​of syntactic structure features from the aforementioned steps. The smart glasses read the preset weight coefficients for these three features from the local configuration, such as W. 时长 W 语义 W 句法 Then, the smart glasses apply a fusion function, such as a linear weighted sum, to calculate the final instruction complexity. The formula is: Instruction Complexity = (Instruction Duration Feature × W) 时长 )+(semantic features × W 语义 )+(syntactic structural features × W 句法 Through this calculation, smart glasses integrate multiple dimensions of linguistic features into a single numerical value that comprehensively reflects the inherent complexity of the instruction; this numerical value is the instruction complexity.

[0048] Step S103: Read the historical dialogue and determine the context complexity score based on the historical dialogue.

[0049] In step S103, historical dialogue refers to the content of one or more dialogue turns preceding the current dialogue turn, stored in the smart glasses' local cache. The context complexity score is a numerical value used to quantify the degree to which the current voice command depends on historical dialogue.

[0050] Specifically, the smart glasses first analyze the instruction text generated in step S102, detecting the presence of pronouns such as "it" or "that," or ellipsis of subjects or objects. If such cases are detected, it indicates that the current voice instruction may be related to historical dialogues. At this point, the smart glasses read the most recent rounds of historical dialogue records from local storage. The smart glasses use a text vectorization model to convert the current voice instruction and each round of historical dialogue into high-dimensional semantic vectors. Semantic relevance is determined by calculating indices such as the cosine similarity between the current instruction semantic vector and the historical dialogue semantic vectors. Simultaneously, the smart glasses count the cumulative number of rounds in the read historical dialogues. Finally, the smart glasses combine the semantic relevance and the cumulative number of rounds to calculate a context complexity score. The higher the semantic relevance or the more rounds of historical dialogue, the higher the context complexity score.

[0051] In one possible implementation, historical dialogues are read, and a context complexity score is determined based on the historical dialogues, specifically including steps S1031-S1034, as follows: Step S1031: Detect whether there are referential words or ellipsis components in the instruction text.

[0052] In step S1031, referential words refer to words in the text that refer to people, things, times, or places that have already appeared in the preceding text, such as "it," "that," or "there." Elliptical components refer to core sentence components such as subjects, predicates, or objects that are omitted in a sentence because the context is already clear.

[0053] Specifically, the natural language processing unit of the smart glasses analyzes the transcribed instruction text. This analysis process includes two aspects: First, the smart glasses scan and match the instruction text in a pre-defined referential lexicon to identify the presence of words such as "he," "she," "it," "this," and "that." Second, the smart glasses perform shallow syntactic analysis on the instruction text to check whether the main structure of the sentence is complete. For example, for the instruction "Another cup," the smart glasses will identify that the instruction lacks an object, thus determining the presence of an ellipsis. If referential words or ellipsis are detected, the smart glasses mark the current instruction as context-dependent.

[0054] Step S1032: When a referential word or ellipsis is detected, read the historical dialogue stored in the local cache and vectorize the historical dialogue to obtain the historical semantic vector.

[0055] In step S1032, local cache refers to a non-volatile storage area on the smart glasses device used for temporary data storage. Historical dialogue refers to one or more rounds of interaction records between the user and the smart glasses that occurred before the current dialogue. Vectorization refers to the process of converting text information into a multi-dimensional numerical vector using a specific text embedding model. Historical semantic vector refers to a numerical vector that represents the core semantics of the historical dialogue, obtained by vectorizing the content of the historical dialogue.

[0056] Specifically, when the detection result in step S1031 is context-dependent, the smart glasses trigger the historical dialogue reading process. The smart glasses access a specially designated dialogue record area in the local cache and read a certain number of recently stored dialogue rounds, such as the last five rounds. Each round of dialogue records includes the user's previous questions and the smart glasses' previous responses. Next, the smart glasses call a text vectorization model deployed on the device side, inputting all the read historical dialogue text content as a whole into the model. After calculation, the model outputs a high-dimensional numerical vector that can summarize the overall semantics of this historical dialogue; this vector is the historical semantic vector.

[0057] Step S1033: Calculate the semantic correlation between the current semantic vector and the historical semantic vector corresponding to the voice command, and count the cumulative number of rounds in the historical dialogue.

[0058] In step S1033, the current semantic vector refers to the numerical vector obtained by processing the voice command text input by the current user using the same vectorization model as the one used to generate the historical semantic vector. Semantic relevance refers to the numerical value obtained by calculating the similarity between the current semantic vector and the historical semantic vector, used to quantify the semantic relevance between the current command and the historical dialogue. The cumulative number of rounds refers to the sum of the interaction rounds contained in the historical dialogue records read from the local cache.

[0059] Specifically, the smart glasses use the same text vectorization model as in step S1032 to convert the instruction text of the current round into a current semantic vector. Then, the smart glasses employ a cosine similarity calculation method to calculate the cosine value of the angle between the current semantic vector and the historical semantic vectors in multidimensional space. This cosine value is the semantic relevance, ranging from -1 to 1; the closer to 1, the more semantically relevant. Simultaneously, the smart glasses count the historical dialogue records read from the local cache, obtaining an integer value, which represents the cumulative number of rounds.

[0060] Step S1034: Determine the context complexity score based on semantic relevance and cumulative round count.

[0061] In step S1034, the context complexity score is a final score that measures the amount of contextual information required to understand the current instruction, which is derived by combining the semantic relevance between the current instruction and the historical dialogue, as well as the length of the historical dialogue.

[0062] Specifically, the smart glasses take the semantic relevance score and the cumulative number of rounds obtained in step S1033 as input. The smart glasses fuse these two scores using a preset formula. For example, this formula could be a weighted summation function: Context Complexity Score = (Semantic Relevance × Weight Coefficient A) + (Logarithmically quantified cumulative number of rounds × Weight Coefficient B). Using this formula, the higher the semantic relevance or the more rounds of historical dialogue, the higher the calculated context complexity score. If no referential words or ellipsis are detected in step S1031, the context complexity score is directly set to 0. The final value obtained is the context complexity score.

[0063] Step S104: Combine the context complexity score and instruction complexity to determine the final execution complexity.

[0064] In step S104, the final execution complexity refers to the comprehensive complexity evaluation value used for the final decision, which combines the complexity of the instruction itself and its dependence on the context.

[0065] Specifically, the smart glasses use the instruction complexity generated in step S102 and the context complexity score generated in step S103 as two input variables. The smart glasses calculate the instruction complexity and context complexity score using a preset fusion function, such as weighted summation. For example, the final execution complexity can be calculated as the instruction complexity multiplied by a weight coefficient W1, plus the context complexity score multiplied by a weight coefficient W2. This calculation process outputs a single numerical value, which is the final execution complexity, comprehensively reflecting the total resource overhead required to process the current voice command.

[0066] Step S104: Based on the final execution complexity, determine the target processing mode that matches the final execution complexity in the preset processing modes.

[0067] In step S104, the final execution complexity refers to a comprehensive value obtained by fusing the instruction complexity and context complexity scores, which quantifies the total resource consumption required to process the current voice command. The preset processing mode refers to a set of predefined operation strategies in the smart glasses system, each strategy corresponding to a different computing resource allocation scheme, such as a purely local processing mode, a cloud-based collaborative processing mode, or a fully cloud-based processing mode. The target processing mode refers to the single processing mode selected from the preset processing mode set that is most suitable for executing the current voice command, based on the specific value of the final execution complexity.

[0068] Specifically, the smart glasses first use a pre-defined fusion algorithm to combine the instruction complexity and context complexity scores calculated in previous steps to generate a final execution complexity. For example, the final execution complexity can be calculated as a weighted sum of the instruction complexity and context complexity scores. Then, the smart glasses compare this final execution complexity value with a pre-defined complexity range threshold table. This threshold table divides the entire complexity range into multiple consecutive intervals, and each interval uniquely maps to a pre-defined processing mode. For example, a final execution complexity of 0 to 30 maps to a local processing mode, 31 to 70 to a cloud-based collaborative processing mode, and 71 and above to a fully cloud-based processing mode. The smart glasses determine the target processing mode that matches the calculated final execution complexity by identifying which interval it falls into.

[0069] In one possible implementation, the preset processing modes include a pure local high-speed mode, a local slicing hybrid mode, and an edge-cloud collaborative mode. Based on the final execution complexity, a target processing mode that matches the final execution complexity is determined from the preset processing modes. Specifically, this includes steps S1041-S1043, as follows: Step S1041: If the final execution complexity is less than or equal to the first preset threshold, then the target processing mode is determined to be the pure local high-speed mode.

[0070] In step S1041, the pure local high-speed mode refers to a preset processing mode that relies entirely on the built-in computing unit and storage resources of the smart glasses to complete instruction processing. This mode has the fastest response speed and does not consume network data. The first preset threshold is a pre-set numerical critical point used to distinguish between low-complexity tasks and medium-complexity tasks.

[0071] Specifically, the smart glasses' decision-making system compares the calculated final execution complexity with a first preset threshold stored in the system configuration. If the final execution complexity is less than or equal to the first preset threshold, it indicates that the current voice command structure is simple, does not depend on context, and the execution action can be completed quickly locally. In this case, the system determines the target processing mode as a pure local high-speed mode. For example, for commands such as adjusting volume, querying time, or turning a local function on or off, the calculated final execution complexity will usually fall within this range.

[0072] Step S1042: If the final execution complexity is greater than the first preset threshold and less than the second preset threshold, then the target processing mode is determined to be the local slice hybrid mode, and the second preset threshold is greater than the first preset threshold.

[0073] In step S1042, the local slice hybrid mode refers to a preset processing mode in which a medium-complexity task is completed locally on the smart glasses by calling different processing units or executing in steps. The second preset threshold is a pre-set numerical threshold greater than the first preset threshold, used to distinguish between medium-complexity tasks and high-complexity tasks.

[0074] Specifically, when the smart glasses determine that the final execution complexity is greater than a first preset threshold but less than a second preset threshold, the system will determine the target processing mode as a local slice hybrid mode. This typically corresponds to instructions that, while they can be completed locally, require multiple steps or the invocation of more complex local resources. For example, to process an instruction to "play rock music from my collection," the smart glasses first need to understand the intent locally, then access the local music database, filter out songs that match the "rock" tag and belong to the "collection" list, and finally invoke the player to play them. The entire process is broken down into multiple task slices, which are completed locally by invoking different resources.

[0075] Step S1043: If the final execution complexity is greater than or equal to the second preset threshold, then the target processing mode is determined to be the edge-cloud collaborative mode.

[0076] In step S1043, the edge-cloud collaboration mode refers to a preset processing mode in which the smart glasses and the remote cloud server work together to complete the instruction processing. This mode is used to process instructions that require massive amounts of data, complex calculations, or real-time information.

[0077] Specifically, if the final execution complexity calculated by the smart glasses is greater than or equal to a second preset threshold, the system determines that local resources are insufficient to independently and efficiently process the current voice command. At this point, the system determines the target processing mode as an edge-cloud collaborative mode. The smart glasses will send the command request, which has undergone preliminary local parsing, along with necessary context information, to the cloud server via network communication. The cloud server utilizes its powerful computing capabilities and vast database to process the command, performing tasks such as complex natural language understanding, large-scale knowledge graph queries, or real-time traffic analysis, and returns the processing results to the smart glasses, which then ultimately present them to the user.

[0078] Step S105: Call the computing resources corresponding to the target processing mode to process the voice command and output the response result.

[0079] In step S105, computing resources refer to a series of hardware and software assets mobilized to execute voice commands, including a central processing unit, memory, network interface, and access permissions to specific algorithm models and databases deployed locally or in the cloud. The response result refers to the feedback information generated and presented to the user by the smart glasses after processing the voice command; this information can be in the form of voice broadcast, text display, or image presentation.

[0080] Specifically, after determining the target processing mode, the smart glasses' resource scheduling system immediately executes that mode. If the target processing mode is local processing, the system will utilize the device's own processor and memory, leveraging lightweight models and data on the edge to complete instruction processing, and output the generated response through the built-in display or speaker. If the target processing mode is cloud-based collaborative processing, the system will package the computationally complex parts of the instruction, such as large-scale data retrieval, and send them to the cloud server via the network. Simultaneously, it will process the remaining parts of the instruction locally. After the cloud returns the results, the smart glasses will integrate the results locally to form a complete response and output it. If the target processing mode is fully cloud-based processing, the smart glasses will encrypt the complete instruction information and necessary context data before sending it to the cloud. The cloud server will complete all processing logic, and the smart glasses will only act as a receiving and presentation terminal, directly presenting the final response returned by the server to the user.

[0081] Please refer to Figure 2 In one possible implementation, the computing resources corresponding to the target processing mode are invoked to process the voice command and output a response result, specifically including steps S201-S207, as follows: Step S201: In pure local high-speed mode, the voice command is input to the dedicated neural processing unit on the smart glasses side for processing, and the dedicated neural processing unit generates the response result. The dedicated neural processing unit is a lightweight model.

[0082] In step S201, the dedicated neural processing unit refers to a low-power coprocessor integrated in the smart glasses hardware, specifically designed for performing neural network calculations. A lightweight model refers to a neural network model that has been pruned and quantized, resulting in a very small number of model parameters and computational load, specifically designed for efficient operation on resource-constrained edge devices.

[0083] Specifically, when the target processing mode is determined to be the pure local high-speed mode, the smart glasses system routes the transcribed voice command text directly to a dedicated neural processing unit. This unit internally stores a lightweight model trained to recognize and execute a limited set of high-frequency, simple commands, such as "turn up the volume," "what time is it," or "turn on the flashlight." Upon receiving the command, the dedicated neural processing unit performs rapid intent matching and execution without going through the main computing unit, generating a corresponding response, such as directly adjusting the system volume or displaying the current time on the screen.

[0084] Step S202: In the local slicing hybrid mode, call the locally deployed full large language model and activate the model slicing engine.

[0085] In step S202, the locally deployed full-scale language model refers to a complete, unedited, large-scale language model with comprehensive natural language understanding and generation capabilities, which is stored intact in the local memory of the smart glasses. The model slicing engine refers to a software or hardware controller responsible for dynamically managing and scheduling different functional layers within the full-scale language model.

[0086] Specifically, when the target processing mode is determined to be the local slicing hybrid mode, the smart glasses' operating system loads the full large language model stored in local flash memory into memory and sends an activation command to the model slicing engine. The model slicing engine then enters the working state, preparing to dynamically allocate power consumption and computing resources to the internal structure of the full large language model according to subsequent instructions.

[0087] Step S203: The model slicing engine activates and powers the first neuron layer corresponding to simple reasoning in the full large language model according to the voice command, and cuts off the power supply to the second neuron layer related to deep reasoning. The first neuron layer includes a understanding layer and a simple logic layer, and the second neuron layer includes a reasoning layer and a generation layer.

[0088] In step S203, the first neuron layer refers to the front-end layer in the full-scale large language model network structure, responsible for performing basic semantic understanding and simple logical judgments. The second neuron layer refers to the back-end deep layer in the model network structure, responsible for complex logical reasoning, association, and creative text generation. The understanding layer is used to parse vocabulary and sentence structure, the simple logic layer is used to handle direct conditions and relationships, the reasoning layer is used to perform multi-step inference and causal analysis, and the generation layer is used to organize and produce fluent and complex natural language.

[0089] Specifically, the model slicing engine first performs rapid pre-analysis of the voice commands to determine their complexity. For moderately complex commands, the engine precisely supplies power to the hardware computing array that constitutes the first neuron layer, activating the understanding layer and the simple logic layer. Simultaneously, the engine uses power gating technology to completely cut off power to the inference and generation layers that constitute the second neuron layer, rendering these high-power neural networks inactive and significantly reducing the overall power consumption of the model.

[0090] In one possible implementation, the model slicing engine activates and powers the first neuron layer corresponding to simple reasoning in the full large language model according to the voice command, and cuts off the power supply to the second neuron layer related to deep reasoning, specifically including steps S2031-S2035, as follows: Step S2031: Pre-store the complete parameters of the full large language model in the non-volatile memory of the smart glasses, and establish a hierarchical index table of model parameters. The hierarchical index table includes the starting storage address, parameter size and inter-layer dependency relationship of each transformer layer.

[0091] In step S2031, non-volatile memory refers to a storage device that can permanently retain data even when power is off, such as the flash memory chip built into smart glasses. Complete parameters refer to all weights and biases required to construct the full large language model. The hierarchical index table is a data structure that records the physical storage information and logical relationships of each component within the full large language model. The transformer layer refers to the basic structural unit that constitutes the backbone network of the modern large language model. The starting storage address, parameter size, and inter-layer dependencies represent the specific storage location of data in the non-volatile memory for each transformer layer, the amount of storage space it occupies, and the order in which data flows during model computation, respectively.

[0092] Specifically, during the production or firmware update phase of the smart glasses, a complete, unmodified full language model is burned into the smart glasses' non-volatile memory. To enable subsequent rapid slice loading, the system parses the model at this stage, creating an entry for each transformer layer and recording the starting address, byte size, and dependency information of the layer's weight parameters in the non-volatile memory, including its layer number and the layer from which its computational input originates. All these entries together form a hierarchical index table, which is also stored in the non-volatile memory as part of the model assets.

[0093] Step S2032: Analyze the task type of the voice command, query the preset task level mapping table, determine the first neuron layer group and the second neuron layer group, and generate the slice activation mask.

[0094] In step S2032, the preset task level mapping table refers to a predefined lookup table that establishes a correspondence between different types of user tasks and the range of transformer levels to be invoked. The slice activation mask is a sequence of binary bits, each bit of which corresponds to a transformer layer in the full language model and is used to indicate whether the layer needs to be activated.

[0095] Specifically, when the model slicing engine receives a voice command, it first performs a quick intent analysis to determine which predefined task type the command belongs to, such as "simple question and answer," "local device control," or "information query." Then, the model slicing engine uses this task type as a keyword to search in a pre-defined task hierarchy mapping table. This table returns the transformer layers that need to be activated to perform this type of task; for example, "simple question and answer" might only require activation of layers 1 to 6. These layers that are determined to need activation collectively form the first neuron layer group, while the remaining unspecified layers in the model form the second neuron layer group. Finally, the model slicing engine generates a slice activation mask, the length of which is equal to the total number of layers in the model, marking 1 at the positions corresponding to layers that need to be activated and 0 at the positions corresponding to layers that need to be disabled.

[0096] Step S2033: Based on the slice activation mask and the layer index table, read the weight parameters corresponding to the first neuron layer group from the non-volatile memory through the memory access controller, and move the weight parameters to the buffer of the on-chip static random access memory.

[0097] In step S2033, the memory access controller refers to a hardware unit specifically responsible for efficiently transferring data between different storage devices, typically a direct memory access controller (DMA). Weight parameters are the values ​​necessary for model calculation. On-chip static random access memory (SRAM) refers to a high-speed, low-latency storage unit integrated within the main processor chip. A buffer is a region within the on-chip SRAM used for temporarily storing weight parameters.

[0098] Specifically, the model slicing engine submits the slice activation mask generated in step S2032 and the hierarchical index table established in step S2031 to the memory access controller. The memory access controller traverses the slice activation mask. For each bit marked as 1 in the mask, the memory access controller looks up the starting storage address and parameter size of the weight parameters of that layer in the hierarchical index table according to the layer corresponding to that bit. Subsequently, the memory access controller initiates a direct memory access operation, bypassing the main computing unit, directly reading all weight parameter data blocks of that layer from the low-speed non-volatile memory, and quickly transferring them to the designated buffer of the high-speed on-chip static random access memory tightly coupled with the neural network computing array, preparing the data for the upcoming computation.

[0099] Step S2034: Send a power consumption control signal to the hardware-level power management unit, apply the operating voltage to the multiply-accumulate array corresponding to the first neuron layer group and turn on the clock signal.

[0100] In step S2034, the hardware-level power management unit refers to a dedicated hardware circuit responsible for managing and distributing power to different functional units within the chip. The power consumption control signal refers to the electrical signal used to instruct the hardware-level power management unit to perform specific power operations. The multiply-accumulate array refers to a parallel computing unit in a neural network accelerator that performs core matrix operations and consists of a large number of multipliers and adders.

[0101] Specifically, simultaneously with or after the weight parameters are moved to the on-chip static random access memory, the model slicing engine sends a precise power control signal to the hardware-level power management unit (HMU). This signal indicates which physical computation units need to be activated. Upon receiving the signal, the HMU connects the power supply to the multiply-accumulate arrays corresponding to the first neuron layer, applying normal operating voltages to these arrays. At the same time, the clock generator begins providing a synchronization clock signal to these powered-on arrays, waking them from their dormant state and putting them into a ready-to-work state.

[0102] Step S2035: The clock tree of the corresponding computing unit of the second neuron layer is turned off by clock gating technology, and the power supply rail of the corresponding computing unit of the second neuron layer is cut off by voltage domain isolation technology, so that the second neuron layer is in a zero static power consumption state.

[0103] In step S2035, clock gating is a design technique that reduces dynamic power consumption by shutting down the clock signal in a specific circuit region. A clock tree refers to the wiring network in a chip that distributes clock signals from the source to various computing units. Voltage domain isolation is a power management technique that eliminates static leakage current by physically switching off the power supply to a specific circuit region. A computing unit power rail refers to the physical wire that provides power to the computing unit. Zero static power consumption refers to the ideal state where the computing unit consumes almost no energy when completely powered off.

[0104] Specifically, for those layers marked as 0 in the slice activation mask, i.e., the second neuron layer groups, the model slicing engine sends disable commands to the hardware-level power management unit and the clock controller. First, the clock controller applies clock gating technology, inserting logic gates on the corresponding branches of the clock tree to prevent the clock signal from propagating to the multiply-accumulate array corresponding to the second neuron layer group. Next, the hardware-level power management unit applies voltage domain isolation technology, activating the power switches located on the power rails to physically cut off the power supply to the computing units corresponding to the second neuron layer group. Through the combined effect of these two technologies, the computing units of the second neuron layer group have neither clock drive nor power supply, thus entering a deep sleep state that generates neither dynamic nor static power consumption, i.e., a zero static power consumption state.

[0105] Step S204: The first neuron layer processes the voice command and generates a response result.

[0106] Specifically, after the model slicing engine completes the precise activation and deactivation operations on the neuron layers, the text data of the voice command is fed into the activated first neuron layer. The understanding layer first parses the intent and entities of the command; for example, in the command "Play songs I've saved," it parses key information such as "play" and "saved." Subsequently, the simple logic layer performs a filtering operation in the local music database based on this information. After processing, a portion of the units in the first neuron layer organizes and generates a simple response, such as "Okay, playing for you," or directly executes the playback action.

[0107] Step S205: In the edge-cloud collaborative mode, invoke the intent extraction capability of the locally deployed full-scale large language model to compress the voice command into compressed prompt words containing core semantic information.

[0108] In step S205, intent extraction capability refers to the specific function in the full-scale language model used to quickly analyze user input and grasp its core purpose and key parameters. Compressed prompts refer to a refined and encoded, structured, and concise data string that accurately represents the complete intent of the original speech command with minimal information.

[0109] Specifically, when the target processing mode is determined to be the edge-cloud collaborative mode, the smart glasses briefly activate the understanding layer of the locally deployed full-scale language model. This part of the model processes and converts the original, potentially lengthy, voice commands, such as "Find me the route from here to the square. I want to take the subway, and I want to get there as quickly as possible," into compressed prompts. This process only utilizes the model's shallow understanding capabilities, avoiding the high time consumption and high power consumption associated with deep processing.

[0110] Step S206: Upload the compressed prompt to the cloud server for processing, and while waiting for the cloud to return the result, put the local main computing unit of the smart glasses into a deep standby state.

[0111] In step S206, the cloud server refers to a remote computing center with powerful computing capabilities and massive data resources. The local main computing unit refers to the central processing unit in the smart glasses responsible for main calculations and control. Deep standby state refers to an extremely low-power sleep mode in which most functions of the main computing unit are turned off, retaining only the most basic wake-up logic.

[0112] Specifically, the smart glasses send the compressed prompt generated in step S205 to a designated cloud server via the wireless network module. Once the data transmission is complete, to conserve power as much as possible, the smart glasses' power management system immediately switches the local main computing unit to deep standby mode. In this state, except for the communication chip that maintains the network connection to receive returned data, most other chips are in a sleep state with no or very low power consumption.

[0113] Step S207: Receive and decode the data returned by the cloud server to form a response result.

[0114] Specifically, when the cloud server completes complex calculations and returns the results, the smart glasses' network communication module receives the data and sends an interrupt signal to wake up the local main computing unit, which is in deep standby mode. Once awakened, the local main computing unit reads the received data packets, decodes and parses them, and converts the structured data returned from the cloud into a format that the user can understand. For example, it renders route planning data into a map interface and navigation instructions, or synthesizes complex question-and-answer results into fluent speech, ultimately presenting a complete response to the user.

[0115] In one possible implementation, after calling the computing resources corresponding to the target processing mode to process the voice command and outputting the response result, the method further includes steps S106-S109, as follows: Step S106: Determine whether there is a semantic relationship between the next round of dialogue and the current dialogue.

[0116] In step S106, the next round of dialogue refers to a new voice command initiated by the user immediately after the smart glasses complete an interaction and output a response. The current dialogue refers to a complete interactive round that has just ended, consisting of the user's voice command and the system's response. Semantic relevance means that the next round of dialogue has a logical continuity or relevance to the current dialogue in terms of topic, intent, or referent.

[0117] Specifically, after the smart glasses output their response to the current dialogue, the system does not immediately clear the context information of this interaction. The system enters a brief waiting state and performs rapid semantic analysis on the received voice commands for the next round of dialogue. This analysis process compares the keywords, entities, or intent vectors extracted from the next round of dialogue with the context information retained in the current dialogue, calculating a relevance score. If this score exceeds a preset threshold, the system determines that the next round of dialogue is semantically related to the current dialogue.

[0118] Step S107: If it is determined that the next round of dialogue has a semantic relationship with the current dialogue, then the activated slice of the current dialogue is retained in the cache memory.

[0119] In step S107, the activated slice refers to the first neuron layer group loaded and powered from the full large language model to process the current dialogue. The cache memory refers to a storage space specifically allocated for temporarily storing model slices. This space can be part of non-volatile memory, designed to accelerate subsequent slice calls and avoid repeated readings from the original model file.

[0120] Specifically, once semantic association is determined in step S106, the system performs a caching operation. The system completely copies the currently active slice, whose weight parameters are already stored in on-chip static random access memory, to a designated cache memory area. Simultaneously, the system appends an identifier to this cached slice, such as a label related to the current dialogue task type, for quick retrieval later. The purpose of this retention action is to directly reuse these ready model layers when processing highly relevant next rounds of dialogue, thereby skipping the time-consuming process of reloading from non-volatile memory.

[0121] Step S108: Determine the difference between the target slice and the activated slice required for the next round of dialogue, and load the difference increment into the on-chip static random access memory.

[0122] In step S108, the target slice refers to the new neuron layer group that needs to be activated after analyzing the task type of the next round of dialogue. The difference set refers to the part of the neuron layer that is included in the target slice but not in the activated slice. Incremental loading means loading only the missing neuron layer parameters, rather than loading the entire target slice, thereby reducing data transfer and loading latency.

[0123] Specifically, the system first parses the next round of dialogue to determine the target slice needed to process it. Then, the system compares the hierarchical structure of the target slice with the hierarchical structure of the activated slices cached in the cache memory. Through this comparison, the system identifies neuron layers that exist in the target slice but not in the activated slices; these layers constitute the difference set. Subsequently, the system initiates a loading operation only for these neuron layers in the difference set. Through the memory access controller, it reads the weight parameters corresponding to the difference set from the full model parameters in non-volatile memory and moves these weight parameters to on-chip static random access memory (SRAM). These weight parameters are then merged with the reusable portions of the activated slices to form the complete computational resources serving the next round of dialogue.

[0124] Step S109: Maintain reuse statistics for each slice in the cache memory, and remove slices that have not been reused in a preset number of rounds of dialogue from the cache memory according to the preset minimum usage strategy.

[0125] In step S109, reuse statistics refer to data such as the number of times each cached slice is reused or the timestamp of its most recent use. The preset least-used strategy is a cache eviction algorithm, such as the Least Recently Used (LRU) algorithm, which is used to decide which data to discard when cache space is insufficient. The preset number of rounds of dialogue is a time or event counter that defines the criteria for determining whether a slice is cold data.

[0126] Specifically, the smart glasses system maintains a metadata record for each slice in the cache memory. Whenever a cached slice is successfully reused in a subsequent conversation, the system updates the slice's reuse statistics, such as increasing its usage count or updating its last usage time. The system periodically, or when cache memory space is low, initiates a cleanup process. This cleanup process iterates through the reuse statistics of all cached slices and, based on a preset least-use strategy, identifies slices that have not been called again in a preset number of conversation rounds. These low-value, cold data slices are deleted from the cache memory, freeing up space to cache new, more likely-to-be-reused slices.

[0127] The following describes a resource collaborative scheduling device for smart glasses from the perspective of hardware processing in an embodiment of this invention. Please refer to [link / reference]. Figure 3 This is a schematic diagram of the structure of a resource collaborative scheduling device for smart glasses in an embodiment of this application.

[0128] It should be noted that, Figure 3 The structure of the resource coordination scheduling device for smart glasses shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0129] like Figure 3 As shown, a resource coordination and scheduling device for smart glasses includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage section 308 into a random access memory (RAM) 303, such as executing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for device operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0130] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0131] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.

[0132] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0134] Specifically, the resource collaborative scheduling device for smart glasses in this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements the resource collaborative scheduling method for smart glasses provided in the above embodiment.

[0135] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the resource coordination scheduling device for smart glasses described in the above embodiments; or it may exist independently and not assembled into the resource coordination scheduling device for smart glasses. The storage medium carries one or more computer programs, which, when executed by a processor of the resource coordination scheduling device for smart glasses, cause the resource coordination scheduling device for smart glasses to implement the resource coordination scheduling method for smart glasses provided in the above embodiments.

Claims

1. A resource collaborative scheduling method for smart glasses, characterized in that, The method includes: Get the voice commands entered by the user in the current conversation turn; The linguistic features of the voice command are analyzed to generate the command complexity. The linguistic features include command duration, keywords, and syntactic structure. Read the historical dialogue and determine the context complexity score based on the historical dialogue; The final execution complexity is determined by combining the context complexity score and the instruction complexity. Based on the final execution complexity, a target processing mode that matches the final execution complexity is determined in the preset processing mode; The computing resources corresponding to the target processing mode are invoked to process the voice command and output a response result.

2. The method according to claim 1, characterized in that, The analysis of the linguistic features of the voice command to generate command complexity specifically includes: The voice command is transcribed to obtain the command text, and the command duration corresponding to the command text is extracted. It is then determined whether the command duration is less than a preset duration threshold. Keyword matching is performed on the instruction text to extract semantic features that indicate whether there are words that require external knowledge reasoning; The instruction text is parsed syntactically to extract syntactic structural features that characterize the nesting depth and number of components of the sentence; The instruction duration, semantic features, and syntactic structure features are fused according to preset weights to generate the instruction complexity.

3. The method according to claim 2, characterized in that, The process of reading historical dialogues and determining a context complexity score based on those dialogues specifically includes: Detect whether there are referential words or omitted components in the instruction text; When the referential words or ellipsis are detected, the historical dialogue stored in the local cache is read and the historical dialogue is vectorized to obtain the historical semantic vector. Calculate the semantic correlation degree between the current semantic vector corresponding to the voice command and the historical semantic vector, and count the cumulative number of rounds of the historical dialogue; The context complexity score is determined based on the semantic relevance and the cumulative number of rounds.

4. The method according to claim 1, characterized in that, The preset processing modes include a pure local high-speed mode, a local slicing hybrid mode, and an edge-cloud collaborative mode. Based on the final execution complexity, determining a target processing mode that matches the final execution complexity within the preset processing modes specifically includes: If the final execution complexity is less than or equal to the first preset threshold, then the target processing mode is determined to be a pure local high-speed mode. If the final execution complexity is greater than the first preset threshold and less than the second preset threshold, then the target processing mode is determined to be a local slice hybrid mode, and the second preset threshold is greater than the first preset threshold. If the final execution complexity is greater than or equal to the second preset threshold, then the target processing mode is determined to be the edge-cloud collaborative mode.

5. The method according to claim 4, characterized in that, The process of calling the computing resources corresponding to the target processing mode to process the voice command and output a response result specifically includes: In the pure local high-speed mode, the voice command is input to a dedicated neural processing unit on the smart glasses side for processing, and the dedicated neural processing unit generates the response result. The dedicated neural processing unit is a lightweight model. In the local slicing hybrid mode, the full local language model is invoked and the model slicing engine is activated; The model slicing engine activates and powers the first neuron layer corresponding to simple reasoning in the full-scale large language model according to the voice command, and cuts off the power supply to the second neuron layer related to deep reasoning. The first neuron layer includes a understanding layer and a simple logic layer, and the second neuron layer includes a reasoning layer and a generation layer. The first neuron layer processes the voice command to generate the response result; In the aforementioned end-cloud collaborative mode, the intent extraction capability of the locally deployed full-scale large language model is invoked to compress the voice command into compressed prompt words containing core semantic information. The compressed prompt is uploaded to the cloud server for processing, and while waiting for the cloud to return the result, the local main computing unit of the smart glasses is placed in a deep standby state. Receive and decode the data returned by the cloud server to form the response result.

6. The method according to claim 5, characterized in that, The step of the model slicing engine activating and powering the first neuron layer corresponding to simple reasoning in the full-scale large language model according to the voice command, and cutting off the power supply to the second neuron layer related to deep reasoning, specifically includes: The complete parameters of the full large language model are pre-stored in the non-volatile memory of the smart glasses, and a hierarchical index table of the model parameters is established. The hierarchical index table includes the starting storage address, parameter size and inter-layer dependency relationship of each transformer layer. The task type of the voice command is parsed, and a preset task level mapping table is queried to determine the first neuron layer group and the second neuron layer group, and a slice activation mask is generated; Based on the slice activation mask and the hierarchical index table, the weight parameters corresponding to the first neuron layer group are read from the non-volatile memory through the memory access controller, and the weight parameters are moved to the buffer of the on-chip static random access memory. Send a power consumption control signal to the hardware-level power management unit, apply a working voltage to the multiply-accumulate operation array corresponding to the first neuron layer group and turn on the clock signal; The clock tree of the corresponding computing unit of the second neuron layer is turned off by clock gating technology, and the power supply rail of the corresponding computing unit of the second neuron layer is cut off by voltage domain isolation technology, so that the second neuron layer is in a zero static power consumption state.

7. The method according to claim 6, characterized in that, After the method calls the computing resources corresponding to the target processing mode to process the voice command and outputs a response result, the method further includes: Determine whether the next round of dialogue has a semantic relationship with the current dialogue; If it is determined that the next round of dialogue has a semantic relationship with the current dialogue, then the activated slice of the current dialogue is retained in the cache memory; Determine the difference set between the target slice required for the next round of dialogue and the activated slice, and load the difference set increment into the on-chip static random access memory; Maintain reuse statistics for each slice in the cache memory, and remove slices that have not been reused in a preset number of rounds of dialogue from the cache memory according to a preset minimum usage strategy.

8. A resource collaborative scheduling device for smart glasses, characterized in that, The resource coordination scheduling device of the smart glasses includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the resource coordination scheduling device of the smart glasses to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the resource coordination scheduling device of the smart glasses, the resource coordination scheduling device of the smart glasses performs the method as described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program runs on the resource coordination scheduling device of the smart glasses, the resource coordination scheduling device of the smart glasses performs the method as described in any one of claims 1-7.