AI model lightweight optimization method and device suitable for resource constraint equipment

By constructing a hardware resource adaptation feature set and customizing lightweight optimization processing, the problems of computing power overload, memory overflow and power consumption exceeding the standard of AI models on the RK3588 device were solved, realizing efficient, real-time and accurate intelligent processing of meeting data on the RK3588 device.

CN121835930AActive Publication Date: 2026-04-10XIAMEN RGBLINK SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing AI models perform poorly on the RK3588 device, failing to meet the actual intelligent processing needs of meeting data. They suffer from problems such as computing power overload, memory overflow, and excessive power consumption. Furthermore, existing lightweight technologies fail to accurately match the characteristics of hardware resources, resulting in low operating efficiency, poor data flow, and generated results with missing information and poor scenario adaptability.

Method used

Based on the computing power, memory, and power consumption hardware parameters of the RK3588 device, a hardware resource adaptation feature set is constructed. Data cleaning, format standardization, and feature extraction are performed to generate a low-redundancy meeting data sample set. Based on the hardware feature set, the initial AI model is customized and lightweight optimized by pruning redundant network layers, quantizing model parameters, and compressing model size to generate a lightweight AI meeting summary model adapted to the hardware resource characteristics of RK3588. Model inference and result verification are then performed.

Benefits of technology

It enables stable and efficient operation of AI models on RK3588 devices, ensuring reasonable consumption of computing power, memory and power consumption, ensuring the real-time and accuracy of data processing, and generating accurate, complete and scenario-adaptable results, thereby improving the practicality and scalability of intelligent conferencing processing on the edge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121835930A_ABST
    Figure CN121835930A_ABST
Patent Text Reader

Abstract

The invention discloses an AI model lightweight optimization method and device suitable for a resource constraint device, which are applied to the technical field of data processing, and aims at computing power, memory and power consumption constraints of an embedded device RK3588, the method comprises the following steps: firstly extracting a device resource threshold and constructing a hardware adaptive feature set, and meanwhile, collecting conference voice and text original data; a sample set for model training reasoning is generated through preprocessing; based on a hardware adaptive feature set, performing customized lightweight optimization on an initial AI conference summary model, creating a lightweight model adaptive to RK3588, deploying the lightweight model to equipment, importing the lightweight model into a sample set for reasoning to generate a conference summary, performing verification and correction to obtain summary data meeting requirements, and finally integrating a whole-process link to obtain a conference summary. And an end-side conference data intelligent processing scheme adaptive to the equipment is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a lightweight optimization method and apparatus for AI models suitable for resource-constrained devices. Background Technology

[0002] The following issues exist in existing technologies, resulting in poor performance of AI models on the RK3588 device and making it difficult to meet the intelligent processing needs of actual meeting data: The RK3588 device faces triple resource constraints: computing power, memory, and power consumption. Its hardware parameters, such as NPU computing power, CPU frequency, physical memory capacity, and power handling capacity, all have clearly defined upper limits. General AI conference summary models are not customized for the hardware characteristics of this device. The model's complex network structure, large parameter scale, and large file size lead to problems such as computing overload, memory overflow, and excessive power consumption after deployment, making it unable to complete inference operations stably and efficiently. The audio and text data in conference scenarios exhibit multi-source heterogeneity, including real-time audio streams, multi-format text, and other types, and contains a large amount of redundant noise. Existing data preprocessing methods are not adapted to the resource constraints of the RK3588. Preprocessing algorithms are computationally intensive and complex, not only consuming excessive device resources but also failing to generate low-redundancy, highly adaptable model training and inference sample sets, thus affecting the accuracy and efficiency of subsequent model inference.

[0003] Existing AI model lightweighting techniques are mostly general-purpose optimization solutions that lack a dynamic correlation mechanism between hardware resource constraints and model optimization dimensions. They simply perform network layer pruning, parameter quantization, or size compression, failing to accurately match the hardware resource characteristics of the RK3588. Optimized models may suffer from overly simplified structures leading to compromised meeting summary functionality, or insufficient optimization failing to adapt to device resources, making it difficult to achieve a balance between resource constraints and summary accuracy. In existing technologies, hardware resource analysis, meeting data preprocessing, model lightweighting optimization, edge deployment and inference, and summary result verification are independent processes lacking a systematic collaborative design. Parameter settings and workflow connections at each stage do not fully consider the dynamic resource allocation characteristics of the RK3588 and the real-time requirements of meeting summaries, resulting in low overall efficiency, poor data flow, and potentially incomplete, semantically inconsistent, and scenario-adaptable meeting summary results that fail to meet practical application needs. Summary of the Invention

[0004] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A lightweight optimization method for AI models suitable for resource-constrained devices includes: extracting resource constraint threshold data for the computing power, memory, and power consumption hardware parameters of the embedded device RK3588, and constructing a hardware resource adaptation feature set; collecting raw voice and text data from a meeting scenario, performing data cleaning, format standardization, and feature extraction preprocessing to generate a meeting data sample set suitable for AI model training and inference; and based on the hardware resource adaptation feature set, performing customized lightweight optimization on the initial AI meeting summary model, pruning redundant network layers, quantizing model parameters, compressing model size, and adapting to RK3588. Based on the hardware resource characteristics of the RK3588, a lightweight AI meeting summary model is generated to adapt to the hardware resource characteristics of the RK3588. The lightweight and optimized AI meeting summary model is deployed to the RK3588 device, and a pre-processed meeting data sample set is imported for model inference to generate meeting summary results. The meeting summary results output by the model are validated, and the results are corrected and optimized in combination with the actual meeting content to generate automatic meeting summary data that meets actual needs. The lightweight model optimization process and the meeting data processing inference link are integrated to form an intelligent processing solution for edge meeting data adapted to the RK3588.

[0005] A lightweight optimization device for AI models suitable for resource-constrained devices is provided. The device is used to execute executable instructions to perform the aforementioned lightweight optimization method for AI models suitable for resource-constrained devices.

[0006] Its beneficial effects are as follows: This invention provides a lightweight optimization method for AI models suitable for resource-constrained devices, and constructs an intelligent processing solution for edge-side conference data to address the computing power, memory, and power consumption constraints of the embedded device RK3588. First, a hardware resource parsing algorithm is used to extract device resource thresholds and construct a hardware adaptation feature set. Then, conference voice and text data are collected, and a low-redundancy sample set is generated through preprocessing such as cleaning and standardization. Based on the hardware feature set, a lightweight AI conference summary model is customized by pruning redundant network layers, quantizing parameters, and compressing volume through model-hardware matching logic and a two-layer optimization architecture. After deployment, the sample set is imported for inference, and an accurate summary is generated through four-dimensional verification and targeted correction. Finally, the entire process is integrated to form a collaborative solution.

[0007] This invention precisely adapts to hardware resources, using customized optimization to make the model compatible with RK3588 features, avoiding issues such as computing overload and memory overflow, and ensuring stable operation. It improves data processing efficiency; lightweight preprocessing algorithms and optimized models reduce resource consumption, keeping inference latency within a reasonable range to meet real-time requirements. It ensures summary quality; a four-dimensional verification and correction mechanism ensures accurate, complete, and scenario-appropriate results, maintaining an accuracy rate of over 90%. It achieves end-to-end collaboration, integrating various stages into a standardized solution, adapting to different meeting scenarios, and improving the practicality and scalability of intelligent processing for edge meetings. Attached Figure Description

[0008] Figure 1 A flowchart illustrating a lightweight optimization method for AI models applicable to resource-constrained devices, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a module for a lightweight optimization device for AI models suitable for resource-constrained devices, provided in an embodiment of the present invention. Detailed Implementation

[0009] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. In one embodiment, this application also proposes a lightweight optimization method for AI models suitable for resource-constrained devices.

[0010] In this application embodiment, a lightweight optimization method for AI models suitable for resource-constrained devices is provided, such as... Figure 1 As shown: S101 extracts device resource constraint threshold data for the computing power, memory, and power consumption hardware parameters of the embedded device RK3588, and constructs a hardware resource adaptation feature set.

[0011] In one implementation, resource feature analysis is performed on the computing power, memory, and power consumption hardware parameters of the embedded device RK3588. First, a hardware resource analysis algorithm is used to extract and quantify the device's basic hardware parameters across all dimensions. Then, based on the analysis results, resource constraint thresholds are defined, ultimately constructing a hardware resource adaptation feature set suitable for the development and deployment of edge AI models. The comprehensive extraction of basic hardware parameters relies on the hardware resource analysis algorithm to perform low-level reading and precise acquisition of the RK3588's physical hardware parameters. For the computing power dimension, the floating-point operation capability of the device's NPU, the core operation frequency of the CPU, and the peak computing power data of multi-core collaborative operation are extracted. Simultaneously, the actual operation speed under single-threaded and multi-threaded conditions is collected, for example, to extract the RK3588's NPU computing power of 6 TOPS and the single-core operation frequency of the eight-core CPU of 2.2 GHz. For the memory dimension, the total physical memory capacity of the device, the upper limit of memory that can be allocated to the AI ​​model, memory read / write speed, and memory usage are extracted. The dynamic fluctuation range of the occupied memory is extracted, such as the core memory parameters such as the device's 8GB physical memory, the maximum memory that can be allocated to the AI ​​model of 4GB, and the memory read / write speed of 2133MHz. In terms of power consumption, the power consumption values ​​of the device under different operating states such as full load, light load, and standby are extracted, as well as the device's power consumption protection threshold and the upper limit of power allocation when the AI ​​model is running. For example, the core power consumption parameters such as the RK3588's full load power consumption of 15W, light load power consumption of 5W, and the upper limit of power allocation for AI model operation of 8W are extracted, so as to achieve full coverage extraction of the basic parameters of the three core hardware dimensions of the device: computing power, memory, and power consumption.

[0012] The quantitative definition of hardware resource constraint thresholds is based on extracted basic hardware parameters. Through in-depth decomposition and computational analysis using hardware resource analysis algorithms, and combined with the resource requirements of the edge AI model deployment, specific constraint thresholds are defined for computing power, memory, and power consumption. For computing power, considering the computing power consumption patterns of AI model inference and training, a computing power carrying capacity threshold is defined, clarifying the maximum computing power occupancy ratio during model operation and the computing power threshold for continuous operation. For example, based on the extracted 6 TOPS NPU computing power, the computing power carrying capacity threshold usable by the AI ​​model is defined as 4.8 TOPS, and the computing power occupancy ratio during model operation should not exceed 80% to avoid computing power overload causing device lag. For memory, considering the memory requirements of the AI ​​model's network structure and parameter scale, an upper limit for memory usage is defined, and a warning threshold for dynamic memory usage is set. For example, based on the extracted 6 TOPS NPU computing power, the maximum computing power occupancy ratio for the AI ​​model is defined as 4.8 TOPS, and the computing power occupancy ratio during model operation should not exceed 80% to avoid computing power overload causing device lag. With 4GB of allocable model memory, a memory usage limit of 3.2GB is set, and a dynamic memory usage warning threshold of 2.8GB is set to prevent memory overflow from affecting model operation. Regarding power consumption, power control standards are defined based on the device's power handling capacity and actual heat dissipation characteristics. The upper limit of power consumption and the allowable range of power fluctuations for model operation are clearly defined. For example, for an AI model based on 8W, the power consumption limit is set as follows: the power consumption of the model operation shall not exceed 7.2W, and the power fluctuation range shall not exceed ±1W. This avoids excessive power consumption that may cause the device to overheat or trigger power protection. The three core resource constraint thresholds of computing power capacity threshold, memory usage limit, and power control standards are quantitatively set. At the same time, baseline rules for AI model hardware adaptation and parameter optimization threshold ranges for edge deployment are generated to provide clear hardware constraint standards for subsequent lightweight model optimization.

[0013] The construction of the hardware resource adaptation feature set integrates defined threshold data for computing power, memory, and power consumption resources. Combining the hardware characteristics of the RK3588 with the deployment requirements of the edge AI conference summary model, various data types are characterized. Computing power capacity thresholds, memory usage limits, and power consumption control standards are used as core feature dimensions. Derivative features, such as parameter fluctuation characteristics, dynamic allocation capabilities of hardware resources, and hardware resource adaptation capabilities under different operating states, are also incorporated. All feature data are structured and integrated according to the optimization requirements of the edge AI model, forming the hardware resource adaptation feature set. The core and derived features of this feature set are interconnected, directly connecting to the model-hardware matching logic for subsequent lightweight optimization of the AI ​​model. It serves as the core basis for the customized lightweight optimization of the initial AI conference summary model. Subsequent optimization operations, such as network layer pruning, parameter quantization, and size compression, are all carried out based on this feature set, ensuring a high degree of adaptation between the optimized model and the hardware resource characteristics of the RK3588.

[0014] S102 collects raw voice and text data from the meeting, performs data cleaning, format standardization, and feature extraction preprocessing to generate a meeting data sample set suitable for AI model training and inference.

[0015] In one implementation, the resource constraints of the RK3588 and the low-redundancy data requirements of the edge model are used as the acquisition criteria. Targeted acquisition of voice and text data in meeting scenarios is conducted. The core is to achieve accurate acquisition of effective meeting data through a multi-source data targeted acquisition algorithm, while adding scene labels to the data to achieve feature-based classification of heterogeneous data. Inputs include the computing power, memory, and power consumption resource constraints of the RK3588, as well as the low-redundancy data judgment criteria for lightweight training and inference of the edge AI meeting summary model, including data redundancy rate thresholds, upper limits of data feature dimensions, and the requirement for the proportion of effective information in a single data point. Simultaneously, data source ports such as audio and video acquisition terminals and meeting document storage terminals in the meeting scenario are accessed. Real-time voice stream data and multi-format raw text data are collected for the meeting scenario. The real-time voice stream data includes continuous voice signals from meeting speeches and voice interaction data from multi-person dialogues. The multi-format raw text data includes meeting minutes drafts, PPT text content, chat box interactive text, meeting agenda documents, etc., covering common text formats such as doc, txt, pdf, and ppt. The multi-source data targeted acquisition algorithm performs preliminary screening of the collected raw data based on a preset low redundancy standard, eliminating meaningless blank data and duplicate data. At the same time, it adds scene tags to all valid collected data according to the characteristics of the meeting scenario. The tag dimensions include meeting type, speaking session, discussion topic, and data type, thereby realizing the scenario-based classification of heterogeneous data.

[0016] For enterprise departmental weekly meetings, real-time audio stream data throughout the meeting is collected, along with various text data in different formats, including meeting agenda docx documents, meeting discussion PPTs, and group chat interaction TXT records. A multi-source data targeted acquisition algorithm is used to remove silent segments of audio data and duplicate document copies. The remaining data is then tagged with scenario labels such as "Departmental Weekly Meeting - Work Report - Business Progress - Audio" and "Departmental Weekly Meeting - Problem Discussion - Project Difficulties - PDF Text," ultimately forming multi-source heterogeneous raw data material with scenario tags. The output consists of multi-dimensional scenario-tagged real-time audio stream raw data and multi-format text raw data, free of redundant information after initial filtering, and with scenario-based feature classification completed.

[0017] For the collected raw data, edge-adapted preprocessing operations are sequentially executed, including redundant noise data cleaning algorithms, multi-format data standardization and normalization algorithms, and lightweight semantic feature extraction algorithms. The entire process is tailored to the low-computing-power and low-memory requirements of the RK3588. All preprocessing algorithms employ a lightweight architecture with no complex computational steps, ultimately achieving the goals of removing invalid information, standardizing data formats, and mining core semantic features. The redundant noise data cleaning algorithm sets differentiated cleaning rules for different characteristics of speech and text data. Speech data focuses on noise removal, while text data focuses on deleting redundant information. The algorithm's computational load is adapted to the RK3588's computing power capacity, without large-scale matrix operations. Noise filtering is performed on the real-time conference audio stream data, removing environmental noise, equipment current noise, and silent segments during speech breaks. It also removes repeated speech segments and meaningless interjections, retaining valid speech content and ensuring the continuity and validity of the audio data. Redundant information is removed from the original text data in various formats. Meaningless content such as blank characters, repeated paragraphs, pop-up advertisement text, and irrelevant watermarks are deleted from the text. At the same time, special symbols and garbled characters in the text are processed uniformly to ensure the cleanliness of the text data.

[0018] The multi-format data standardization and normalization algorithm sets unified standardization rules for different types and formats of meeting data, converting heterogeneous data into a unified format that can be recognized by the edge AI meeting summary model. It also normalizes data units and lengths to adapt to the memory storage requirements of the RK3588, avoiding reduced model inference efficiency due to data format chaos and differences in data units. The cleaned, effective speech stream data is uniformly converted into an audio format with a fixed sampling rate and fixed bit depth. Simultaneously, the speech data is framed, with uniform frame lengths and frame shifts set, converting continuous speech streams into equal-length speech frames for convenient batch processing by the model. The volume amplitude of the speech data is normalized, mapping the amplitude range to a fixed interval to eliminate the impact of volume differences on model training. The cleaned text data of different formats were uniformly converted into plain text format. At the same time, the text was processed by word segmentation and part-of-speech tagging. The text encoding format was uniformly set to UTF-8. The length of the text data was normalized and an upper limit was set for the number of characters in a single text data. Text exceeding the upper limit was reasonably segmented, and text below the upper limit was properly padded to ensure the uniformity of text data specifications.

[0019] Unify the converted voice stream data of the weekly regular meeting after cleaning into the WAV format with a sampling rate of 16 kHz and a bit depth of 16 bits. Perform frame segmentation processing with a frame length of 25 ms and a frame shift of 10 ms, and normalize the volume amplitude to the interval [-1, 1]. Unify the text data of the weekly regular meeting in doc, pdf, and ppt formats into the pure text format encoded in UTF-8. Split the text into a sequence of words through word segmentation processing. Set the upper limit of the number of characters in a single text data to 512. Segment the business progress report text exceeding 512 characters, and perform compliance completion on the problem discussion text with less than 512 characters to complete the standardization and normalization of multi-format data. Finally, output the meeting voice frame data and text word segmentation data with unified format, consistent dimension, and standardized specifications.

[0020] Adopt a lightweight semantic feature extraction algorithm, abandon the complex deep feature extraction architecture, and adopt a lightweight natural language processing algorithm adapted to the edge side to meet the business requirements of the meeting summary. Mine the core semantic features from the standardized voice and text data. The feature dimension adapts to the training and inference requirements of the edge side AI model, without high-dimensional complex features, to avoid increasing the memory occupancy and computing power consumption of the RK3588. First, convert the standardized voice frame data into Mel cepstral coefficients features, and then combine the business logic of the meeting summary to extract core semantic features such as speaker features, keyword features, and semantic emotion features in the voice data. The feature dimension is controlled within the low-dimensional range to adapt to the memory requirements of the RK3588.

[0021] Based on the bag-of-words model or TF-IDF algorithm, combined with the business requirements of the meeting summary, extract core semantic features such as meeting topics, discussion points, decision results, and action arrangements from the standardized text word segmentation data.剔除无意义的词汇特征,保留与会议总结强相关的特征信息,同时对特征进行降维处理,保证特征维度适配端侧模型的轻量化要求。对标准化后的周例会语音帧数据,提取梅尔倒谱系数特征后,进一步挖掘出发言者“部门经理”“项目负责人”的身份特征、“业务进度”“项目延期”“解决方案”的关键词特征;对标准化后的周例会文本分词数据,通过TF-IDF算法提取“Q3业务完成率”“项目整改措施”“下周工作安排”等核心语义特征,剔除“的”“了”等无意义词汇特征,将特征维度控制在64维,完成轻量型语义特征提取。最终输出带核心语义特征的、端侧适配型的会议语音特征数据、文本特征数据,特征维度低、与会议总结业务强相关。

[0022] Based on the lightweight adaptation standards of edge AI models corresponding to the hardware resource characteristics of RK3588, a resource compatibility compliance verification algorithm is used to perform a full-dimensional verification of the preprocessed meeting data. The core verification data focuses on the compatibility of the data with the computing power and memory of RK3588, as well as its compatibility with the lightweight training and inference of the edge AI meeting summary model. Valid meeting data is ultimately selected, while data that does not meet the operational requirements of the edge device is removed. The computing power capacity threshold, memory usage limit, and power consumption control standards of RK3588, as well as the low computing power and low memory requirements for lightweight training and inference of the edge AI meeting summary model, are used as core verification standards and input into the resource compatibility compliance verification algorithm. Specific verification indicators include the upper limit of computing power consumption per batch of data during model training, the upper limit of memory usage for data storage, and the upper limit of data feature dimensions. The resource compatibility compliance verification algorithm matches the pre-processed meeting data against preset verification standards one by one. It judges compliance from four dimensions: data computing power consumption, memory usage, feature dimensions, and effective information ratio. Data that meets all four dimensions is considered valid, while data that does not meet any dimension is considered invalid. The algorithm uses lightweight matching logic to adapt to the computing power requirements of RK3588.

[0023] The computational power consumption of the preprocessed data during edge AI model training and inference is calculated to determine if it is below the RK3588's computational power capacity threshold, preventing excessive computational power consumption of a single data point from causing device lag. The storage volume and feature matrix size of the preprocessed data are also calculated to determine if they are below the RK3588's memory usage limit, preventing memory overflow due to excessive data memory usage. The semantic feature dimensions of the preprocessed data are checked to determine if they meet the lightweight feature dimension requirements of the edge AI model, avoiding high-dimensional features that increase the complexity of model training and inference. The proportion of valid meeting information in the preprocessed data is checked to determine if it meets the low-redundancy data requirement, preventing excessive invalid information from affecting the model's training and inference accuracy. For example, the following validation standards were used to validate the preprocessed weekly meeting data: the AI ​​model computing power threshold of 4.8 TOPS, the memory usage limit of 3.2GB, the model feature dimension limit of 128 dimensions, and the effective information content limit of 80%. Long audio segments exceeding 4.8 TOPS, extremely large text data exceeding 3.2GB of memory usage, complex semantic feature data with 256 dimensions, and casual text data with only 70% effective information content were removed. Meeting data meeting all validation standards were retained, completing the resource compatibility compliance validation. The final output consisted of effective meeting audio and text feature data that met the RK3588 hardware resource constraints and the lightweight training and inference requirements of the edge AI meeting summary model.

[0024] For the verified valid meeting data, a lightweight feature integration algorithm was implemented, and a dedicated dataset for the edge device was constructed. Data was layered and lightweightly packaged according to the edge device resource allocation requirements of the RK3588, ultimately generating a meeting data sample set suitable for AI model training and inference. The dataset fully conforms to the operational requirements of the edge device and can be directly imported into the lightweight AI meeting summary model deployed on the RK3588 for training and inference. The lightweight feature integration algorithm fuses and integrates the verified valid meeting audio and text feature data according to meeting scene labels and core semantic feature types, eliminating duplicate and similar features while retaining core effective features. Simultaneously, the integrated features are lightweight compressed to further reduce the storage volume and computational complexity of the feature data, adapting to the memory and computing power requirements of the RK3588. Data is grouped according to meeting scene labels, and data with the same meeting scene and the same core semantic feature type undergo feature fusion. Corresponding features are established between audio and text feature data to achieve feature complementarity between audio and text data, while ensuring that the integrated feature data does not lose the core information required for meeting summarization.

[0025] The valid weekly meeting data after verification is grouped according to the scenario labels "Work Report", "Problem Discussion" and "Next Week's Arrangement". The voice feature data and text feature data in the "Work Report" group are merged and associated with the voice keyword feature and text numerical feature of "Q3 Business Completion Rate". The duplicate "Department Name" feature in the two groups of data is removed. The integrated features are then compressed to reduce the storage volume of the feature data, thus completing the lightweight feature integration. Finally, the lightweight integrated meeting feature data with scenario grouping, complementary features and no redundancy is output.

[0026] Based on the resource allocation requirements of the RK3588, a dedicated meeting dataset for the edge is constructed. This dataset employs a lightweight architecture, free from complex dataset indexes and redundant additional information. Its storage format is compatible with the RK3588's file system, allowing direct reading and retrieval by the edge AI meeting summary model. According to the different needs of AI model training and inference data, the lightweight, integrated meeting feature data is hierarchically divided. Furthermore, based on the RK3588's computing power and memory allocation ratio, the batch size and data volume ratio of training and inference data are set. The training data emphasizes diversity and comprehensiveness, covering different meeting scenarios and information, while the inference data emphasizes lightweightness and efficiency, adapting to the real-time inference requirements of the edge model. For example, based on the resource allocation requirements of the RK3588, the integrated lightweight feature data of the weekly meeting is divided into training data and inference data in a 7:3 ratio. The training data covers meeting data in all scenarios such as "work report", "problem discussion" and "next week's arrangements". The size of a single batch of training data is set to 32 records adapted to the computing power of RK3588. The inference data selects core meeting decisions, action arrangements and other information. The size of a single batch of inference data is set to 16 records adapted to the computing power of RK3588, thus completing the data layering.

[0027] The layered datasets are encapsulated using a lightweight, edge-adapted encapsulation format. During encapsulation, lossless compression is performed to reduce the overall storage size of the datasets. A simple read index is added to the encapsulated datasets to facilitate quick data access and retrieval by the edge AI meeting summary model. The computational complexity of the encapsulation algorithm is adapted to the computing power requirements of the RK3588. The layered weekly meeting training and inference data are encapsulated using a lightweight encapsulation format, with lossless compression to reduce storage size. Simple read indexes categorized by scene labels and data types are added to the encapsulated datasets to ensure that the lightweight AI meeting summary model deployed on the RK3588 can quickly access the data. The final output is a meeting data sample set suitable for AI model training and inference. This sample set is a dedicated lightweight dataset for edge use, meeting the triple resource constraints of computing power, memory, and power consumption of the RK3588, and can be directly imported into the edge AI meeting summary model for training and inference operations.

[0028] S103, based on the hardware resource adaptation feature set, performs customized lightweight optimization on the initial AI meeting summary model, prunes redundant network layers, quantizes model parameters, compresses model size, adapts to the characteristics of RK3588 hardware resources, and generates a lightweight AI meeting summary model adapted to the characteristics of RK3588 hardware resources.

[0029] In one implementation, a hardware resource analysis algorithm is used to perform a multi-dimensional and in-depth breakdown of the computing power capacity threshold, memory usage limit, and power consumption control standards defined for the RK3588. The core is to transform hardware resource constraint indicators into baseline rules and parameter threshold ranges that can directly guide the lightweight optimization of AI models, thereby achieving a precise connection between hardware constraints and model optimization. The algorithm input is the core constraint data in the hardware resource adaptation feature set, and the output is the baseline rules and thresholds for AI model optimization deployed on the edge.

[0030] The core input data includes the RK3588's computing power capacity threshold, memory usage limit, power consumption control standards, and derived features such as parameter fluctuation characteristics and dynamic allocation capabilities from the hardware resource adaptation feature set. It also incorporates the operational characteristics of models summarized from edge AI conferences, including computing power consumption patterns during model inference, memory usage patterns for parameter storage, and power consumption patterns during model computation. The hardware resource analysis algorithm deeply decomposes hardware resource constraints from three dimensions: model layer optimization, parameter optimization, and size optimization. These correspond to optimization requirements for the number of model network layers, parameter accuracy / scale, and model file size, respectively. During the decomposition process, the actual needs of edge model deployment are considered, transforming hardware constraint indicators into quantifiable model optimization indicators. Based on the computing power capacity threshold, the maximum number of computable model network layers, the upper limit of computing power consumption for a single layer, and the computing power allocation ratio for parallel model computation are decomposed, clarifying the core quantitative indicators for model layer pruning.

[0031] Based on the memory usage limit, the maximum storage size of model parameters, the adaptation range of parameter accuracy, and the memory usage threshold of intermediate feature maps are decomposed, clarifying the core quantitative indicators for parameter quantization. Based on power consumption control standards, the upper limit of single-batch power consumption for model computation, the power consumption threshold for different levels of model operation, and the correlation coefficient between model size and power consumption are decomposed, clarifying the core quantitative indicators for model size compression. The decomposed quantitative model optimization indicators are structured and integrated to generate AI model hardware adaptation baseline rules and parameter optimization threshold ranges for edge deployment. The hardware adaptation baseline rules define the overall criteria for lightweight model optimization, including the optimization direction, adaptation principles, and hardware constraint bottom line for model layers, parameters, and size; the parameter optimization threshold ranges set specific quantitative threshold ranges for each dimension of optimization, clarifying the upper limit, lower limit, and optimal value of each optimization indicator.

[0032] Taking the RK3588's computing power threshold of 4.8 TOPS, memory usage limit of 3.2GB, and power consumption control standard of 7.2W as an example, the hardware resource analysis algorithm, after in-depth decomposition, generates hardware adaptation baseline rules as follows: "Model layer count adapts to computing power capacity, parameter precision / scale adapts to memory usage, model volume adapts to power consumption control, and all optimization operations do not exceed the hardware constraint bottom line." Simultaneously, it generates parameter optimization threshold ranges, including model network layer count ≤ 32 layers, single-layer computing power consumption ≤ 0.15 TOPS, parameter storage size ≤ 2.56GB, parameter precision adapted to 4-8 bits, model file size ≤ 500MB, and single-batch model computation power consumption ≤ 6.5W, providing a clear quantitative benchmark for subsequent lightweight model optimization. Finally, it outputs the AI ​​model hardware adaptation baseline rules adapted to edge deployment and multi-dimensional parameter optimization threshold ranges, which directly serve as the core hardware benchmark for subsequent lightweight model optimization.

[0033] By connecting the network structure data and hardware resource adaptation feature set of the initial AI meeting summary model through model-hardware matching logic, a lightweight optimization benchmark for the model is constructed. Simultaneously, a customized lightweight adaptation model is introduced. Combined with model compression algorithms and RK3588 hardware resource profiles, a dynamic computation mechanism of demand-constraint-optimization adaptation is established to achieve dynamic matching and accurate calculation of model optimization requirements, hardware resource constraints, and lightweight optimization operations. The initial AI meeting summary model is analyzed from all dimensions, extracting its core network structure data, including the number, type, and hierarchical connection relationships of network layers, the computational power consumption of each layer, the precision, scale, storage method, and parameter update rules of model parameters, the size and encoding format of the model file, and the volume ratio of each module. The model's meeting summary business characteristics are also recorded, including the core layer for semantic feature extraction, key parameters for meeting information inference, and the core module for generating summary results.

[0034] The extracted model network structure data is compared with the hardware resource adaptation feature set and hardware adaptation baseline rules using a model-hardware matching logic. This matching logic adheres to the principle of "hardware constraints as the baseline, model business characteristics as the core, and maximizing optimization effect as the goal." It compares the current state of model layers, parameters, and volume with hardware adaptation requirements one by one, identifying redundant parts of the model that exceed hardware constraint thresholds or do not meet edge deployment requirements. Simultaneously, it retains the core layers, key parameters, and core modules that enable the meeting summary function. Based on the matching results, a lightweight model optimization benchmark is constructed, clarifying the core objects, optimization scope, retained content, and hardware constraint boundaries for lightweight model optimization, thus defining a precise range for subsequent optimization operations.

[0035] By integrating all hardware information of the RK3588, including its computing power, memory, power consumption, resource constraint thresholds, parameter fluctuation characteristics, and dynamic allocation capabilities, and combining this with the operational characteristics of the edge AI model, a hardware resource characteristic profile of the RK3588 is constructed. This profile comprehensively depicts the device's hardware resource status, adaptability, and operational patterns, serving as the core hardware input for a customized lightweight adaptation model. A customized lightweight adaptation model is introduced, specifically designed for lightweight edge AI models. Its core modules include a requirement analysis module, a constraint verification module, an optimization calculation module, and a result feedback module. The hierarchical connection between these modules is as follows: the requirement analysis module receives input and passes it to the constraint verification module; after constraint verification, it is passed to the optimization calculation module; the optimization calculation result is fed back to the result feedback module; and simultaneously, the result feedback module transmits information back to the requirement analysis module, achieving dynamic iteration.

[0036] The initial AI conference summarized the lightweight optimization requirements of the model, the characteristics of the RK3588 hardware resources, the benchmark for lightweight model optimization, hardware adaptation baseline rules, and parameter optimization threshold ranges. Model compression algorithms (including network layer pruning, parameter quantization, and model volume compression algorithms) are integrated into the optimization calculation module of the customized lightweight adaptation model, achieving a deep integration of algorithms and the adaptation model. Based on the module connection and algorithm fusion of the customized lightweight adaptation model, a dynamic calculation mechanism of requirement-constraint-optimization adaptation is established. This mechanism can analyze model optimization requirements in real time, verify hardware resource constraints, calculate specific parameters for lightweight optimization operations, and dynamically adjust the optimization strategy based on the optimization calculation results. If the optimization result exceeds the hardware constraint threshold, the optimization magnitude is adjusted in reverse; if the optimization result does not achieve the expected goal, the calculation parameters are optimized within the hardware constraint range. This achieves dynamic matching and accurate calculation of requirements, constraints, and optimization, ensuring that each optimization operation conforms to hardware constraints and meets model optimization requirements.

[0037] For example, the network structure data of the initial AI conference summary model was extracted as 64 network layers, 32-bit parameter precision, 3.8GB parameter storage size, and 800MB model file size. After matching the model with hardware resources and the RK3588 hardware resource adaptation feature set, 32 redundant network layers, 1.24GB of redundant parameter storage size, and a compressible model size of 300MB were identified. The benchmark for lightweight model optimization was to "prune redundant network layers, quantize parameters to 4-8 bits, compress the model size to within 500MB, and retain the semantic feature extraction layer". The core business modules include "meeting information inference parameters, etc." After introducing a customized lightweight adaptation model, it integrates model compression algorithms such as network layer pruning, parameter quantization, and size compression. The input model optimization requirement is "to achieve lightweighting while meeting the meeting summary function and adapting to RK3588 hardware," along with an RK3588 hardware resource characteristic profile. The established dynamic calculation mechanism can calculate in real time the optimized parameters for pruning the network layer to 32 layers, quantizing parameters to 6 bits, and compressing the model size to 450MB. These parameters, after constraint verification, meet the hardware adaptation threshold requirements. The final output includes the model lightweight optimization benchmark, the RK3588 hardware resource characteristic profile, and the dynamic calculation mechanism for requirement-constraint-optimization adaptation. This dynamic calculation mechanism can directly output the specific lightweight optimization parameters for model layers, parameters, and size.

[0038] Using model optimization dimensions (network layers, model parameters, and model volume) as coverage dimensions, the optimization range corresponding to each dimension is deeply integrated with hardware resource adaptation data. Simultaneously, combined with the generated model hardware adaptation baseline rules and parameter optimization threshold ranges, the integration and calibration of each optimization dimension are completed. This ensures that optimization operations in each dimension conform to hardware constraints and baseline rules, and that optimizations in each dimension are mutually coordinated and conflict-free. Based on the optimization parameters output by the dynamic computation mechanism of demand-constraint-optimization adaptation, the specific optimization ranges for the three optimization dimensions—network layers, model parameters, and model volume—are defined, clarifying the optimization objects, optimization magnitudes, and optimization objectives for each dimension. The core hardware constraints for each dimension's optimization are also labeled to ensure that the optimization range remains within the parameter optimization threshold range. The system clearly defines the redundant network layer numbers to be pruned, the core network layer numbers to be retained, the total number of pruned network layers, and the upper limit of computational power consumption for each layer; it also clearly defines the parameter modules to be quantized, the parameter precision conversion range, the storage scale of the quantized parameters, and the upper limit of parameter memory usage; and it clearly defines the model modules to be compressed, the volume compression ratio, the compressed model file size, and the upper limit of power consumption for model computation.

[0039] The defined optimization dimensions are deeply integrated with hardware resource adaptation data. This data includes core and derived features of the hardware resource adaptation feature set, the hardware operating characteristics of RK3588, and the correlation between model operation and hardware resources. A dimension-based fusion algorithm is employed, using each optimization dimension as the core. This algorithm matches information related to that dimension in the hardware resource adaptation data, adding hardware adaptation tags to each optimization dimension. It clarifies the hardware adaptation requirements, resource allocation ratios, and operational constraints for each optimization operation, achieving precise integration of the optimization scope and hardware resources, ensuring a high degree of model-hardware resource compatibility. The integrated optimization dimension data is then compared and calibrated against the model hardware adaptation baseline rules and parameter optimization threshold ranges. If the integrated optimization dimension data exceeds the parameter optimization threshold range or does not conform to the hardware adaptation baseline rules, the optimization scope is readjusted based on a dynamic calculation mechanism of demand-constraint-optimization adaptation until all optimization dimension data conforms to the baseline rules and is within the parameter optimization threshold range. Simultaneously, the synergy between optimization dimensions is ensured to prevent over-optimization of a single dimension from causing other dimensions to fail to meet hardware constraints.

[0040] Based on the optimization parameters output by the dynamic computation mechanism, the optimization dimensions were defined as follows: network layers were pruned to 32 layers, redundant layers numbered 33-64 were pruned, and core layers numbered 1-32 were retained. Parameters were quantized to 6 bits, the quantized parameter module used all parameters, the quantized storage size was ≤2.56GB, the model size was compressed to 450MB, the compression ratio was 43.75%, and the compression module was the redundant feature extraction module. Through a dimension association fusion algorithm, this optimization range was fused with RK3588 hardware resource adaptation data, and hardware adaptation tags were added to each dimension, such as single-layer computing power consumption ≤0.15TOPS after network layer pruning, memory usage ≤2.56GB after parameter quantization, and single-batch power consumption ≤6.5W after model size compression. The fused data was compared and calibrated with the hardware adaptation baseline rules and parameter optimization threshold range to confirm that all optimization data conformed to the baseline rules and were within the threshold range, thus completing the integration of each optimization dimension. The final output is multi-dimensional optimized and integrated data of network layers, model parameters, and model volume after fusion and calibration. This data provides accurate, compliant, and collaborative optimization basis for subsequent lightweight processing operations.

[0041] Through a lightweight processing mechanism involving network layer pruning, parameter quantization, and volume compression, the initial AI meeting summary model undergoes actual lightweight optimization operations. This integrates model structure and hardware constraint data, while a customized lightweight adaptation model enhances the accuracy of optimization and adaptation. Ultimately, the dual-layer optimization architecture of simplified model structure and efficient parameter compression is combined to generate a lightweight AI meeting summary model adapted to the RK3588 hardware resource characteristics. Based on the multi-dimensional optimized and integrated data after fusion and calibration, a lightweight processing mechanism for network layer pruning, parameter quantization, and volume compression is constructed. This mechanism integrates the three lightweight operations, clarifying the execution order, execution standards, hardware constraint requirements, and business function retention criteria for each operation. A real-time hardware constraint verification module is also introduced to ensure that the operation adheres to hardware resource constraints throughout the entire execution process.

[0042] Based on the scope defined by the optimized and integrated data, a network layer pruning algorithm was used to precisely prune redundant network layers in the initial model, retaining only the core network layers that implement the meeting summary function. During the pruning process, the computational power consumption of each layer was verified in real time to ensure that the total computational power consumption of the pruned network layers was less than or equal to the computational power carrying capacity threshold, while maintaining the hierarchical connection relationship of the core network layers to avoid affecting the model's semantic feature extraction and meeting information reasoning capabilities. After network layer pruning, a parameter quantization algorithm was used to convert the model parameters to higher precision and optimize their scale. According to the parameter precision range defined by the optimized and integrated data, high-precision parameters were converted to low-precision parameters suitable for edge devices, while redundant parameters were eliminated. During quantization, the parameter storage scale was verified in real time to ensure that the memory usage of the quantized parameters was less than or equal to the upper limit of memory usage, while retaining the validity of key parameters.

[0043] After network layer pruning and parameter quantization, a model volume compression algorithm is used to losslessly compress the model file. Redundant model modules and invalid feature data are compressed. During compression, the power consumption of the model is checked in real time to ensure that the power consumption corresponding to the compressed model volume is less than or equal to the power consumption control standard, while ensuring the integrity and operability of the model file. During the execution of the lightweight processing mechanism, the results of each optimization operation are input into the customized lightweight adaptation model in real time. The model's constraint verification module and optimization calculation module perform real-time calculations and verifications. If there is a deviation between the optimization result and the hardware adaptation requirements, the model will output dynamic adjustment instructions to correct the parameters of the lightweight operation in real time, such as adjusting the number of network layer pruning, the accuracy of parameter quantization, and the volume compression ratio, to ensure the accuracy of the optimization operation and achieve a high degree of integration between the model structure and hardware constraint data.

[0044] This design deeply integrates the dual-layer optimization architecture features of model structure simplification and efficient parameter compression. Model structure simplification, the first layer of optimization, achieves lightweight model architecture through network layer pruning to adapt to the computing power constraints of the RK3588. Efficient parameter compression, the second layer of optimization, focuses on lightweighting model parameters and files through parameter quantization and size compression to adapt to the memory and power consumption constraints of the RK3588. The two optimization layers support each other and progress progressively. The first layer lays the structural foundation for the second, while the second layer enables resource adaptation for the first. The integration process ensures the synergy between the two layers, achieving overall model lightweighting and comprehensive hardware resource adaptation. After completing all lightweight optimization operations and integrating the two-layer optimization architecture, a lightweight AI meeting summary model adapted to the hardware resource characteristics of RK3588 was generated. The hardware compatibility of the model was verified, including full-dimensional verification of computing power consumption, memory usage, and power consumption. This ensured that when the model was running on the RK3588 device, all resource consumption was within the hardware constraint thresholds. At the same time, the meeting summary function of the model was verified to ensure that the model still maintains good accuracy and effectiveness in summarizing meeting content after being lightweighted.

[0045] For example, the constructed lightweight processing mechanism explicitly executes in the order of "network layer pruning → parameter quantization → volume compression," with the execution standard being to conform to hardware adaptation baseline rules and operate within the parameter optimization threshold range. The network layer pruning algorithm prunes the initial model's 32 redundant network layers, retaining 32 core layers. Real-time verification confirms that the model's computational power consumption after pruning is 4.2 TOPS ≤ 4.8 TOPS, meeting the computational power carrying capacity threshold. The parameter quantization algorithm quantizes the model parameters from 32 bits to 6 bits, eliminating redundant parameters. Real-time verification confirms that the parameter storage size after quantization is within 2.4GB ≤ 3.2GB. Storage usage limit; the model size was compressed from 800MB to 450MB using a model volume compression algorithm, and real-time verification confirmed that the power consumption of the compressed model in a single batch was 6.0W≤7.2W, meeting the power consumption control standard; the optimization results of each step were input into a customized lightweight adaptation model, and after verification, there was no deviation, and no adjustment of optimization parameters was required; finally, a two-layer optimization architecture of model structure simplification and efficient parameter compression was integrated to generate a lightweight AI meeting summary model. Verification showed that the resource consumption of this model running on the RK3588 met hardware constraints, and the meeting summary accuracy remained above 90%, meeting the needs of practical applications. The final output is a lightweight AI meeting summary model adapted to the hardware resource characteristics of the RK3588. This model achieves comprehensive lightweighting of the network layer, parameters, and size, and is highly adapted to the computing power, memory, and power consumption hardware characteristics of the RK3588, and can be directly deployed to this device for meeting summary inference calculations.

[0046] S104 deploys the lightweight and optimized AI meeting summary model to the RK3588 device, imports the pre-processed meeting data sample set for model inference, and generates meeting summary results.

[0047] In one implementation, the inputs include the deployment goals of the RK3588 edge device, such as model stability requirements, inference latency limits, and resource consumption control targets; the inputs also include AI model inference performance requirements, such as meeting summary accuracy thresholds, semantic integrity standards, and minimum inference speed; and the inputs further include meeting summary scenario adaptation characteristics, such as the summary focus of different meeting types, output format requirements, and user reading habits. The multi-dimensional information integration algorithm employs a three-dimensional mapping logic of "goal-requirement-characteristics," associating the edge device deployment goals, inference performance requirements, and scenario adaptation characteristics with four operation items: model deployment, data import, resource allocation, and result output. Weighted calculations determine the setting range of each operation item, while establishing collaborative constraints between operation items to ensure that parameters at each stage are mutually compatible and conflict-free.

[0048] Define the model's runtime environment configuration on the RK3588, model loading method (e.g., static loading, dynamic loading), and inference engine selection (e.g., TensorFlowLite, ONNXRuntime). Set upper limits for model initialization time and the number of threads in the inference process to ensure the model can start quickly and run stably after deployment. Define the import format (e.g., JSON, binary) of the preprocessed conference data sample set, data transmission rate limits, and batch import size. Set data import verification rules (e.g., data integrity verification, format compliance verification) to avoid affecting the inference process due to data import anomalies. Based on the computing power, memory, and power consumption constraints of the RK3588, define the computing power allocation ratio, memory usage upper limit, and power consumption control range during model inference. Define the peak resource consumption for a single batch of inference to ensure that the inference process does not exceed the device's resource carrying capacity.

[0049] Based on the adaptability characteristics of meeting summary scenarios, the core dimensions of the output results are clearly defined, including meeting topics, core discussion points, decision results, action items, and time nodes. Information integrity requirements and language expression standards are set for each dimension to ensure that the output results align with users' actual needs. Collaborative constraint rules are established between various operation items. Model deployment adaptation parameters must match the inference resource allocation threshold, sample data import specifications must be compatible with the input format of model deployment, and the output dimensions of the summary results must be consistent with the output structure of the AI ​​model inference, ensuring seamless connection and efficient collaboration across all stages of the deployment and inference process. For example, the deployment goals for the RK3588 edge are model inference latency ≤500ms and memory usage ≤2.5GB; the AI ​​model inference performance requirements are meeting summary accuracy ≥90% and semantic completeness ≥85%; and the meeting summary scenario adaptability characteristics are that work meetings need to highlight action items and time nodes, while project review meetings need to highlight decision results and risk points. After processing by a multi-dimensional information integration algorithm, the scope of each operation item is clearly defined: Model deployment adaptation parameters are: using the TensorFlowLite inference engine, static loading method, initialization time ≤100ms, and 2-4 inference threads; Sample data import specifications are: JSON format, batch import size 16 records / batch, and data transfer rate ≤10MB / s; Inference resource allocation thresholds are: computing power allocation ratio ≤70% (≤3.36TOPS), memory usage limit 2.5GB, and power consumption controlled at 5-6W; Summary results output dimensions are: meeting topics, core discussion points, decision results, action items, and time nodes, with information integrity ≥90% for each dimension. Collaboration requirements are: the number of inference threads does not exceed the number of CPU cores of RK3588 (4), the JSON data format must conform to the TensorFlowLite input specifications, and the output dimensions must correspond one-to-one with the model inference output structure.

[0050] This algorithm constructs a complete logical chain for edge-side inference deployment by designing an inference logic algorithm. It clarifies the execution standards for core and auxiliary operations, ensuring efficient use of device resources while guaranteeing the accuracy and effectiveness of meeting summary results. The algorithm aims for both "efficient resource utilization" and "effective inference results," employing a "core-auxiliary" two-layer logical architecture. Core operations focus on the core processes of model deployment and data import, ensuring the stability and reliability of the basic inference steps. Auxiliary operations focus on resource scheduling and feature matching, improving inference efficiency and result accuracy. A dynamic association mechanism enables collaborative operation between core and auxiliary operations. The lightweight AI meeting summary model execution standards include model file integrity verification (ensuring no missing model weights or network structure files), runtime environment compatibility testing (verifying the adaptability of dependent libraries, system versions, and models), inference function connectivity testing (confirming the model can normally receive input data and output results), and resource consumption testing (monitoring computing power, memory, and power consumption during model initialization and inference). During deployment and debugging, all test data must be recorded. If resource consumption exceeds limits or functional abnormalities occur, the model deployment adaptation parameters must be adjusted until the requirements are met.

[0051] The execution standards for targeted import of preprocessed conference data samples include format conversion before import (converting the sample set into a model-readable format), batch data splitting (splitting the sample set according to the set batch size), real-time monitoring of the import process (monitoring data transmission rate and integrity), and inference input format verification (ensuring that the dimensions and types of the imported data are consistent with the model input requirements). Importing and inference should be performed batch by batch, with model inference triggered after each batch of data is imported to avoid overloading equipment resources due to batch imports.

[0052] The execution criteria for dynamic scheduling of computing power during the inference process include real-time monitoring of the computing power occupancy status of the RK3588 (collecting computing power data every 10ms), dynamically adjusting the number of inference threads based on computing power occupancy (reducing one inference thread when computing power occupancy is ≥90%; adding one inference thread when computing power occupancy is ≤50%, with the number of threads not exceeding the set range), and prioritizing the computing power supply for the core inference process (reserving at least 30% of computing power for model inference when the device has other tasks). This dynamic scheduling algorithm enables flexible allocation of computing power resources, ensuring that the inference process is uninterrupted and does not experience lag.

[0053] The execution standard for real-time semantic feature matching of conference data includes extracting semantic features from the imported data (extracting keywords and core semantic vectors based on the TF-IDF algorithm), performing real-time matching with the feature library used during model training (calculating feature similarity), and adjusting inference parameters based on the matching results (if the feature similarity is ≥85%, the default inference parameters are used; if the similarity is between 60% and 85%, the semantic weight parameters are adjusted to enhance matching accuracy; if the similarity is <60%, scene adaptation supplementary parameters are enabled). Through the real-time semantic feature matching algorithm, the model's adaptability to different conference data is improved, ensuring the semantic accuracy of the summary results. For example, in the core operation items, the model deployment and debugging on the client side requires completing model file verification (confirming that the 450MB lightweight model file is complete), runtime environment testing (verifying the compatibility between TensorFlow Lite 2.15 and RK3588), functional connectivity testing (the model can output summary results after inputting test data), and resource usage testing (memory usage of 1.2GB during initialization and computing power usage of 3.0 TOPS during inference); data import and inference requires splitting the JSON format sample set into batches of 16, monitoring the transmission rate to be stable at 8MB / s during import, triggering inference after each batch of data import, and verifying that the data dimension is 64 dimensions consistent with the model input requirements. In the auxiliary operation items, dynamic computing power scheduling requires real-time monitoring of computing power utilization. When the computing power utilization reaches 92% when the number of inference threads is 4, one thread is reduced to 3, at which point the computing power utilization drops to 75%; real-time semantic feature matching requires extracting the keywords of a batch of meeting data, such as "project progress issue rectification next week's arrangement", matching the feature library with a similarity of 78%, adjusting the semantic weight parameters, and then performing inference.

[0054] To address the consistency requirements of summary results across various lightweight meeting processing scenarios, inference optimization rules were established, including dynamic adaptation of deployment parameters, compliance verification of sample import, real-time adjustment of resource allocation, and precise feature matching. Multi-dimensional inference optimization rules were set through rule optimization algorithms to ensure that the model outputs stable and consistent meeting summary results under different meeting scenarios and data quality conditions, while also adapting to the device resource characteristics of the RK3588. With "scenario adaptation" and "result consistency" as its core objectives, the algorithm constructs a dynamically adjusted optimization rule system based on the characteristics of different lightweight meeting scenarios (such as meeting duration, data type, summary focus), device resource fluctuation characteristics, and data quality differences. Rules are prioritized to achieve collaborative operation.

[0055] The model deployment parameters are dynamically adjusted based on the meeting scenario type (e.g., for short, quick meetings, the number of inference threads is adjusted to 2 to reduce power consumption; for long, complex meetings, the number of threads is adjusted to 4 to improve inference speed). The model loading method is also dynamically adjusted based on device resource fluctuations (when memory usage is close to the limit, a dynamic loading method is used to load the model in modules; when memory resources are sufficient, a static loading method is used to improve startup speed). Through a dynamic adaptation algorithm for deployment parameters, real-time matching of model deployment with scenarios and resources is achieved.

[0056] Set data format validation rules (validate the integrity of JSON data fields and the correctness of data types), data quality validation rules (validate data redundancy rate ≤10% and effective information ratio ≥80%), and data size validation rules (the amount of data imported in a single batch shall not exceed the set 16 records). Through compliance validation algorithms, invalid data is eliminated to prevent invalid data from affecting the consistency of inference results, while also reducing unnecessary resource consumption.

[0057] Resource allocation is dynamically adjusted based on inference progress (60% of computing power is allocated for model initialization in the early stages of inference, 70% for core computation in the middle stages, and 50% for result integration in the later stages). Resource allocation is also adjusted based on power consumption (when power consumption exceeds 6.5W, computing power allocation is reduced by 10% to decrease power consumption; when power consumption is below 5W, computing power allocation is increased by 5% to improve efficiency). By adjusting the algorithm in real-time through resource allocation, inference efficiency is ensured while keeping device power consumption within a safe range.

[0058] For different meeting scenarios, feature matching thresholds are set (≥75% for routine work meetings, ≥80% for project review meetings), matching strategies are adjusted based on data quality differences (fast matching algorithms are used for high-quality data, and precise matching algorithms are used for low-quality data), and a feature update mechanism is established (features from new meeting data are regularly added to the feature library to improve matching accuracy). Through precise feature matching and adaptation algorithms, the consistency of the model's understanding of the data in different scenarios is ensured, thereby guaranteeing the consistency of the summary results.

[0059] For example, in the dynamic adaptation rules for deployment parameters, for a 15-minute quick weekly meeting, the number of inference threads is adjusted to 2, and power consumption is controlled at 5.2W; for a 90-minute project review meeting, the number of threads is adjusted to 4, increasing the inference speed by 30%. In the sample import compliance verification rules, when verifying a batch of data, two data entries were found to have a redundancy rate of 18%, which were directly removed, and the data were imported for inference at a rate of 14 entries per batch. In the real-time adjustment rules for resource allocation, 60% of the computing power (2.88 TOPS) is allocated for model initialization in the initial stage of inference, which is increased to 70% (3.36 TOPS) for computation in the middle stage, and then reduced to 50% (2.4 TOPS) for result integration in the later stage, with the power consumption remaining stable at around 5.8W throughout. In the precise feature matching adaptation rules, the feature matching threshold for the project review meeting is set at 80%. The feature matching degree of a batch of data is 79%. After enabling the precise matching algorithm, the matching degree is increased to 82%, ensuring that the summary results meet the key requirements of the project review meeting.

[0060] This algorithm systematically integrates operational requirements, resource utilization and inference result effectiveness coordination needs, and inference optimization rules through a dataset integration algorithm. This generates a standardized baseline dataset for edge-side meeting inference. Based on this dataset, edge-side model inference operations are then performed, ultimately generating the meeting summary results. The algorithm employs a "classification-relationship fusion-standardized encapsulation" process. First, information from each dimension is classified according to deployment type, execution specifications, collaborative logic, and optimization strategies. Then, a correlation algorithm establishes mapping relationships between different categories of information. Finally, the dataset is encapsulated in a standardized format to form a basic dataset, ensuring a clear dataset structure, complete information, and direct model usability.

[0061] Deployment type information includes core information such as model deployment method (static loading), inference engine (TensorFlowLite), and runtime environment configuration, clearly defining the core attributes of model deployment. Execution specification information includes execution standards, parameter ranges, and threshold requirements for each operation, such as data import batch size of 16 records / batch, computing power allocation range of 2.4-3.36 TOPS, and feature matching threshold ≥75%. Collaboration logic information includes collaborative constraints between operations, linkage mechanisms between core and auxiliary operations, and priority ranking of rules, ensuring coordinated operation of all stages of the inference process. Optimization strategy information includes the specific execution logic of optimization rules such as dynamic adaptation of deployment parameters, computing power scheduling, and feature matching, providing a basis for dynamic adjustments to the inference process.

[0062] After model deployment, the standardized edge-side meeting inference dataset is read. Key data such as model deployment parameters, sample data types, and inference resource usage are extracted using data parsing algorithms. Data is then sorted according to its correlation with inference accuracy and meeting summary efficiency (inference accuracy weighted at 60%, summary efficiency weighted at 40%), prioritizing highly correlated data. Based on the sorted data, the model performs inference operations according to the set deployment inference logic, execution standards, and optimization rules. First, semantic understanding and feature extraction are performed on the imported data. Then, a lightweight network layer is used for inference analysis, integrating key information such as meeting topics, core discussion points, and decision results to form a preliminary summary. The preliminary summary is formatted according to the set output dimensions to ensure concise language, clear logic, and complete information. Simultaneously, the results are verified to meet inference performance requirements (accuracy ≥ 90%, semantic completeness ≥ 85%). If the requirements are met, the results are output directly; otherwise, the inference optimization rules are adjusted and the calculation is repeated.

[0063] For example, after dataset integration, the deployment type information is clearly defined as statically loading the TensorFlowLite engine; the execution specifications information includes a data import scale of 16 records / batch and a computing power range of 2.4-3.36 TOPS; the collaboration logic information clarifies the linkage between computing power scheduling and the number of inference threads; and the optimization strategy information includes parameter adjustment logic for scenario adaptation. After the model reads the dataset, it prioritizes processing project review meeting data with high relevance based on relevance. During inference operations, it extracts core information such as project progress risk points and decision opinions from the data. After inference analysis by the network layer, a preliminary summary result is formed. After formatting processing, the accuracy rate reaches 92% and the semantic completeness reaches 88%, meeting the performance requirements. The final output is the meeting summary result.

[0064] S105 performs data verification on the meeting summary results output by the model, and corrects and optimizes the results based on the actual meeting content to generate automatic meeting summary data that meets actual needs.

[0065] In one implementation, the core inputs are the meeting summary results output by the model and a preprocessed meeting data sample set (including actual meeting audio and raw text data). A four-dimensional verification system of "accuracy, completeness, consistency, and scenario adaptation" is constructed through a multi-dimensional result verification algorithm to comprehensively check the quality of the summary results and identify deviations and deficiencies. The meeting summary results output by the model are input, including dimensions such as meeting topics, core discussion points, decision results, and action items. The preprocessed meeting data sample set is input, from which key information such as actual meeting audio transcripts, original meeting documents, and scenario tags are extracted as verification benchmarks. The verification standards include a summary accuracy threshold (≥90%), an information completeness threshold (≥85%), a semantic consistency threshold (≥88%), and a scenario adaptation threshold (≥92%). The multi-dimensional result verification algorithm adopts a three-layer operation logic of "benchmark comparison - feature matching - logical verification". First, the summary results are compared with the core information of the actual meeting data. Then, the semantic relevance is checked through the feature matching algorithm. Finally, the rationality of the content is verified through the logical verification algorithm. Each dimension of verification is independent of each other and the results are weighted and integrated. Finally, a verification report and a list of deviations are output.

[0066] The four-dimensional verification is as follows: For accuracy verification, the core keywords (such as project name, decision conclusion, and time node) of the summary results and the actual meeting data are extracted based on the TF-IDF algorithm, and the keyword matching degree is calculated; the semantic similarity algorithm (such as cosine similarity) is used to compare the semantic fit between the summary results and the key paragraphs of the actual meeting. If the matching degree or fit is lower than the accuracy threshold, it is marked as accuracy deviation.

[0067] For completeness verification: Based on the core dimensions of the meeting summary (topics, discussion points, decisions, action items, etc.), an information missing detection algorithm is used to check whether the summary results omit key information from the actual meeting, such as whether important decisions are not mentioned or core action items are missing. The percentage of missing information is calculated, and if the percentage exceeds 15% (i.e., completeness is below 85%), it is marked as a completeness deviation. For consistency verification: A logical consistency algorithm is used to check whether the internal semantics of the summary results are self-consistent (e.g., no contradictory decision descriptions). At the same time, the timeline of the summary results is compared with the actual meeting data, and whether the participants' statements are consistent. If there are contradictions or inconsistent statements, it is marked as a consistency deviation. For scenario adaptation verification: Combining meeting scenario tags (e.g., work meetings, project review meetings), a scenario feature matching algorithm is used to check whether the summary results fit the summary focus of the scenario (e.g., work meetings highlight action items, review meetings highlight risk points). If the focus deviates or the expression style is inconsistent with the scenario, it is marked as a scenario adaptation deviation.

[0068] For example, the model outputs the summary of a project review meeting, mentioning that "the project will be launched in Q4, with no risk points." Accuracy verification shows that the extracted keyword "project launch Q4 risk points" matches the actual meeting data with 82% (below 90%), and the semantic similarity is 80%. Completeness verification reveals that the three core risk points identified in the actual meeting were not mentioned, resulting in a 30% missing rate. Consistency verification found no contradictory statements. Scenario adaptation verification shows that the focus of the risk assessment in the review meeting was not highlighted, with a scenario adaptation rate of 80%. Finally, a verification report is generated, marking three types of deviations: accuracy, completeness, and scenario adaptation, and listing the specific deviation points.

[0069] Based on the deviations identified in the verification report and using actual meeting data, a deviation correction optimization algorithm is employed to target and correct the summary results. Simultaneously, the expression logic and format are optimized to ensure the corrected results accurately reflect the actual meeting situation. The deviation correction optimization algorithm follows a "deviation classification - targeted correction - logic optimization" process. First, deviations are categorized by type (accuracy, completeness, consistency, scenario adaptation). Then, corresponding correction information is extracted based on actual meeting data to precisely correct the deviations. Finally, the language and logical structure of the summary results are optimized to ensure concise and clear content.

[0070] To address issues of insufficient keyword matching and semantic similarity, precise expressions are extracted from actual meeting data to replace vague or erroneous content in the summary results. A keyword enhancement algorithm incorporates frequently occurring core terms from actual meetings (such as "project phase two iteration" and "core function acceptance") into the summary results, improving accuracy. A semantic calibration algorithm corrects expressions that do not align with the semantics of the actual meetings. To address information gaps, an information completion algorithm extracts missing key information (such as unmentioned decisions, action items, and risk points) from actual meeting data and supplements it according to the dimensions of the summary results, ensuring completeness across all core dimensions. A redundancy filtering algorithm is used during the supplementation process to avoid duplication of new information with existing content. For semantic contradictions or inconsistencies, a logical calibration algorithm compares the summary results with actual meeting data to determine correct expressions and correct inconsistencies. For inconsistencies in timelines, participants, and other information, a unified calibration is performed based on the actual meeting records to ensure consistency between the summary results and the actual meeting data. To address the issue of scenario-focused deviation, a scenario feature enhancement algorithm is used to highlight the core summary elements of the corresponding scenario (e.g., work meetings emphasize action items and time nodes, project review meetings emphasize risk points and decision-making basis); a style adaptation algorithm is used to adjust the language style of the summary results (e.g., formal meetings use standardized expressions, quick meetings use concise expressions) to fit the needs of the scenario.

[0071] Based on the verification report from the previous project review meeting, regarding the accuracy deviation correction, "the project will be launched in Q4" was revised to "the second phase iteration of the project will be launched after the core functions are accepted in Q4 2024," supplementing the accurate statement from the actual meeting; regarding the completeness deviation correction, three core risk points ("technical architecture compatibility risk," "third-party interface integration delay risk," and "insufficient testing resources risk") were extracted from the actual meeting data and added to the summary results in the format of "risk point - countermeasures"; regarding the scenario adaptation deviation correction, the summary structure was adjusted to "project progress - core decisions - risk assessment - follow-up plans," emphasizing the risk assessment focus of the review meeting, and the language style was adjusted to formal and standardized expression.

[0072] The revised summary results undergo a second verification to ensure all deviations have been eliminated. Simultaneously, the quality of the summary is further improved through optimized iterative algorithms, ultimately generating automated meeting summary data that meets actual needs. Using the same multi-dimensional verification system and algorithm as the initial verification, the revised summary results are comprehensively checked, focusing on verifying whether the original deviations have been corrected and whether any new deviations have been added. If the score for any dimension still does not reach the threshold during the second verification, the process returns to the previous step for re-correction; if all dimensions meet the standards, the process enters the optimization iteration phase.

[0073] The optimized iterative algorithm focuses on the practicality and readability of the summarized results. It uses a concise expression algorithm to eliminate redundant expressions and merge repetitive semantics, compressing the text length (while keeping the core information unchanged); a logical sorting algorithm to optimize the arrangement order of content in each dimension (such as sorting by the meeting process of "topic-discussion-decision-action item") to improve clarity; and a keyword highlighting algorithm to strengthen the expression of key information such as core terms, time nodes, and responsible parties (such as clearly stating the person in charge of the action item and the deadline), thereby improving reading efficiency.

[0074] S106 integrates the model lightweight optimization process with the conference data processing inference link to form an intelligent processing solution for end-side conference data that is compatible with RK3588.

[0075] In one implementation, the entire process of summarizing the AI-powered meeting on the edge is meticulously broken down, clarifying the boundaries and core operations of five core process links. Simultaneously, three types of results are constructed using process decomposition algorithms and data flow tracking algorithms: process link segmentation results, link interaction association tables, and data flow mapping relationships. This achieves visualization and standardization of the entire process operation, links, and data. The core of the hardware resource analysis process is to extract the computing power, memory, and power consumption hardware parameters of the RK3588 from all dimensions using hardware resource parsing algorithms, quantify and define resource constraint thresholds, and ultimately construct a hardware resource adaptation feature set to provide hardware constraint benchmarks for all subsequent processes. The core of the meeting data preprocessing link is to complete the collection, cleaning, standardization, feature extraction, and compliance verification of raw meeting audio and text data using a series of lightweight algorithms, such as multi-source data targeted acquisition algorithms and redundant noise data cleaning algorithms, ultimately generating a meeting data sample set adapted for edge model training and inference.

[0076] The core of the lightweight model optimization process involves using dimensional correlation fusion algorithms and model-hardware matching logic to connect the hardware resource adaptation feature set with the network structure data of the initial AI meeting summary model. This process sequentially completes network layer pruning, parameter quantization, and model size compression, ultimately generating a lightweight AI meeting summary model adapted to RK3588. The core of the edge-side deployment and inference chain involves using multi-dimensional information integration algorithms and inference logic design algorithms to complete the deployment and configuration of the lightweight model on RK3588, targeted import of sample sets, dynamic allocation of inference resources, and model inference operations, ultimately generating preliminary meeting summary results. The core of the summary result verification and optimization chain involves using multi-dimensional result verification algorithms and deviation correction optimization algorithms to construct a four-dimensional verification system to check the quality of the preliminary summary results. Targeted correction and secondary verification are performed for deviations, ultimately generating automatic meeting summary data that meets actual needs.

[0077] The document clearly defines the core operational steps, execution entities, algorithm support, and hardware constraints of the five major process links. For example, the hardware resource analysis process is divided into "hardware parameter extraction → resource threshold quantification → feature set construction," with the hardware resource parsing module as the execution entity, the hardware resource parsing algorithm as the algorithm support, and the computing power constraint not exceeding the RK3588 computing power carrying capacity threshold of 4.8 TOPS. It also clarifies the sequential relationships, triggering conditions, and coordination requirements of each operational step within the five major process links. For instance, after the "feature set construction" step of the hardware resource analysis process is completed, the "model-hardware matching logic operation" step of the model lightweight optimization process is triggered, and the feature set data must be synchronized to the model lightweight optimization module in real time, with the data transmission format being a dedicated structured format for the edge side. Track the generation, transmission, processing and output paths of various core data in the entire process, and clarify the data format, scale, transmission rate and processing algorithm. For example, "the hardware resource adaptation feature set is generated by the hardware resource analysis process, transmitted to the model lightweight optimization process in 64-dimensional feature vector format, with a transmission rate of no more than 10MB / s, and after being processed by the model-hardware matching logic algorithm, it is transformed into quantitative index data for model lightweight optimization".

[0078] Based on the triple constraints of computing power, memory, and power consumption of the RK3588, and the low redundancy and high real-time requirements of lightweight conference processing, a process standardization algorithm was used to optimize and adjust the three types of core mapping results generated in the early stage. Simultaneously, the granularity of process integration was set according to hardware resource adaptation characteristics, and the link linkage response cycle was clarified based on the real-time requirements of the conference summary, ensuring that all process links conform to the operating characteristics of the end-side devices and business needs. Complex operation links in each process link that were incompatible with the RK3588 hardware constraints were eliminated and replaced with lightweight operations. For example, the "feature matching link of large-scale matrix operation" in the model lightweight optimization process was replaced with a "lightweight feature matching algorithm link adapted to the end-side," reducing computing power consumption. At the same time, the execution standards of each process link were unified. For example, the "lower limit of effective information ratio" for all data processing links was uniformly set to 80%, consistent with the low redundancy requirements of conference data preprocessing.

[0079] Simplify non-core interactions between different stages and reduce unnecessary data exchanges across modules. For example, remove unnecessary triggering conditions between the "model lightweight optimization process" and the "summary result verification optimization process," retaining only the core connection of "sending a start signal to the edge-side inference link after model optimization is completed." Simultaneously, clearly define the hardware resource usage limits for interactions between each stage; for example, the memory usage for cross-stage data transfer should not exceed the RK3588 memory limit of 3.2GB. Unify the data transmission format across the entire process to JSON format adapted to the RK3588 file system, compressing the data transmission scale. For example, convert model parameter data from 32-bit floating-point format to 6-bit quantization format before transmission. Furthermore, optimize data flow paths and reduce data transfer steps. For example, directly transfer the conference data sample set from the "conference data processing module" to the "edge-side model inference module" to avoid inference delays caused by third-party module transfers.

[0080] The granularity of process integration is set in conjunction with the hardware resource adaptation features of RK3588. Based on the core criteria of "hardware resource constraint adaptability," "operational link synergy," and "data processing efficiency," the granularity of process integration is set at the "sub-process / sub-link level." This means that the integration operation only merges the five major process links as a whole, without breaking down the core operational links within each process link, ensuring the integrity and independence of each process. For example, based on RK3588's dynamic computing power allocation capability, the "hardware resource analysis process" is integrated into a "hardware resource adaptation sub-process," retaining its complete internal steps of "parameter extraction, threshold quantization, and feature set construction" without further breakdown, avoiding fragmented consumption of computing power due to excessively fine granularity.

[0081] The defined response cycle for each workflow link is based on the real-time requirements of lightweight meeting processing. Combined with the computing speed of the RK3588, a real-time requirement calculation algorithm is used to determine the response cycle between each workflow link, keeping the overall response cycle within 500ms to meet the real-time inference requirements of the edge side. Specifically, the response cycle for core workflow links is set to a shorter threshold; for example, "after the edge-side inference link completes inference and generates preliminary summary results, the verification step of the summary result verification and optimization link must be triggered within 100ms." The response cycle for non-core workflow links can be appropriately relaxed; for example, "after the hardware resource adaptation sub-process completes feature set updates, the updated data must be synchronized to the lightweight model construction sub-process within 300ms." Simultaneously, the set response cycle must match the computing power and speed of the RK3588 to avoid device lag due to excessively short cycles.

[0082] Based on established process integration granularity standards and link linkage response cycles, a process grouping and integration algorithm is used to group the five major process links in relation to each other. Process links with tightly coupled operational logic and frequent data flow are merged into a unified core system, ultimately forming four core end-side conference processing systems: hardware resource adaptation sub-process, conference data processing sub-link, lightweight model construction sub-process, and edge-side inference optimization sub-link. This achieves intensive management of process links and improves end-side processing efficiency. The hardware resource adaptation sub-process directly integrates all operational steps of the original "hardware resource analysis process," serving as the core hardware foundation system for the entire end-side conference processing. It provides real-time hardware resource constraint benchmarks and adaptation feature data for all other core systems. The core algorithm of this system is a hardware resource parsing algorithm, which can monitor the computing power, memory, and power consumption status of the RK3588 in real time, dynamically update the hardware resource adaptation feature set, and synchronize the updated feature data to related systems such as the lightweight model construction sub-process and the edge-side inference optimization sub-link according to the established link linkage response cycle.

[0083] The conference data processing sub-link fully integrates all operational steps of the original "conference data preprocessing link," serving as the core data source system for edge-side conference processing. It provides training data for the lightweight model construction sub-process and inference data for the edge-side inference optimization sub-link. This system relies on a series of edge-adaptive algorithms, including multi-source data targeted acquisition algorithms, redundant noise data cleaning algorithms, and lightweight semantic feature extraction algorithms, to complete the entire preprocessing of heterogeneous raw conference data from multiple sources. The generated conference data sample set will be transmitted to the lightweight model construction sub-process and the edge-side inference optimization sub-link respectively, according to a unified transmission format and linkage response cycle. Real-time compliance verification is performed during data transmission, ensuring that the data meets the resource constraints of RK3588.

[0084] The lightweight model construction sub-process integrates all operational steps of the original "lightweight model optimization process." As the core system of the model for edge-side meeting processing, it is a crucial link connecting hardware resources and data processing. Based on the hardware resource adaptation feature set provided by the hardware resource adaptation sub-process, this system completes customized lightweight optimization of the initial AI meeting summary model through model-hardware matching logic, dimension association fusion algorithms, and network layer pruning algorithms. The generated lightweight AI meeting summary model is directly transmitted to the edge-side inference optimization sub-link, and hardware compatibility verification is performed during model transmission to ensure that the model's computing power, memory, and power consumption are all within the RK3588 constraint thresholds.

[0085] The edge-side inference optimization sub-link integrates all operational steps of the original "edge-side deployment inference link" and "summary result verification optimization link," serving as the final execution core system for edge-side meeting processing and a crucial step in achieving meeting summary result output. This system incorporates core algorithms such as multi-dimensional information integration algorithms, inference logic design algorithms, and multi-dimensional result verification algorithms. It first completes the deployment and inference computation of the lightweight model on the RK3588, generating preliminary meeting summary results. Then, it performs full-dimensional verification and targeted correction on the results, ultimately outputting automatically generated meeting summary data that meets actual needs. Simultaneously, the system feeds back resource consumption data and result verification quality data during the inference process to the hardware resource adaptation sub-process and the lightweight model construction sub-process in real time, providing a basis for hardware feature set updates and model iterative optimization.

[0086] The four core edge conferencing systems are prioritized for execution. Then, combined with the RK3588's dynamic resource allocation rules and the edge performance requirements of the AI ​​model, a full-process collaborative operation plan is formulated through a full-process collaborative algorithm. Finally, a complete edge conferencing data intelligent processing solution is generated, which includes system number, process composition details, execution timing plan, resource adaptation rules, and scenario adaptation requirements. This achieves efficient collaboration of the four core systems, reasonable resource allocation, and precise adaptation to conferencing scenarios.

[0087] The execution priority of the four core systems is based on the principle of "hardware foundation as a prerequisite, data as support, model as the core, and inference output as the goal." Considering the edge-side operating characteristics of the RK3588, the execution priority of the four core systems is determined as follows: Hardware resource adaptation sub-process (highest priority) > Meeting data processing sub-link > Lightweight model construction sub-process > Edge-side inference optimization sub-link. The highest priority hardware resource adaptation sub-process must be started first, completing hardware parameter extraction and feature set construction in real time to provide hardware constraint benchmarks for the operation of all subsequent systems. The meeting data processing sub-link and the lightweight model construction sub-process can run in parallel to improve overall processing efficiency. The edge-side inference optimization sub-link, as the final execution stage, must be started after both the model and data are ready to ensure smooth inference operations.

[0088] Based on the execution priority and computing power, memory, and power consumption characteristics of the four core systems, differentiated resource allocation ratios are set for each system. For example, 5% of computing power and memory are allocated to the hardware resource adaptation sub-process to ensure the stability of its real-time monitoring; the highest resource ratio (computing power ≤ 70%, memory ≤ 2.5GB) is allocated to the edge-side inference optimization sub-link to meet the core requirements of inference and verification; and approximately 20%-30% of resources are allocated to the conference data processing sub-link and the lightweight model construction sub-process to support their parallel operation. Simultaneously, a dynamic resource scheduling mechanism is set up so that when a system is under high load, the resource allocation ratio can be temporarily adjusted without affecting the core operation of other systems, adapting to the dynamic operating characteristics of the devices.

[0089] Based on the core performance requirements of the model, namely "summary accuracy ≥90%, inference latency ≤500ms, and semantic integrity ≥85%", execution standards and collaboration requirements are set for each system. For example, the effective information ratio of the sample set generated by the conference data processing sub-link must be ≥80% to ensure the quality of the data source for model training and inference; the lightweight model construction sub-process must ensure that the summary accuracy of the lightweight model does not decrease significantly and the inference speed is improved by more than 30%; the edge inference optimization sub-link must strictly control the inference latency to ensure that the total time from data import to result output is ≤1s to meet the real-time requirements of the edge.

[0090] A complete intelligent processing solution for edge-side conference data is generated based on priority ranking and a full-process collaborative operation scheme. This solution, adapted to the RK3588, comprises five core elements, each implemented using algorithms and tailored to hardware and business requirements. A unique edge-side operation number is assigned to each of the four core systems: H01 (hardware resource adaptation sub-process), D02 (conference data processing sub-link), M03 (model lightweight construction sub-process), and R04 (edge-side inference optimization sub-link), enabling precise identification and management of each system. The entire operational process, core supporting algorithms, execution standards, and hardware constraints of each core system are clearly defined. For example, the detailed process of the R04 edge-side inference optimization sub-link is: "Model deployment and debugging → data-oriented import → inference calculation → result verification → deviation correction → result output." The core algorithms are a multi-dimensional information integration algorithm and a multi-dimensional result verification algorithm. The execution standards are a summary accuracy ≥90% and a computing power constraint ≤3.36 TOPS.

[0091] Based on priority ranking and linkage response cycles, a timeline plan for the startup, operation, coordination, and termination of each core system is formulated. The startup time difference, parallel operation window, data transmission time nodes, and result output time thresholds for each system are clearly defined. For example, H01 starts first, D02 and M03 start in parallel 50ms later, and R04 starts within 100ms after M03 completes model optimization. The total execution time of R04 is ≤800ms. The allocation ratios and dynamic adjustment rules for computing power, memory, and power consumption of each core system under different operational stages and meeting scenarios are clearly defined. For example, in long project review meetings, 10% more computing power is allocated to the inference phase of R04; in short, fast-paced weekly meetings, the proportion of computing power allocated to model optimization in M03 is reduced, and power consumption control requirements are increased.

[0092] Based on the summary needs of different meeting scenarios (work meetings, project review meetings, thematic seminars, etc.), scenario-based execution requirements are set for each core system. For example, for project review meetings, D02 needs to focus on extracting core semantic features such as "project risks, decision results, and rectification measures," and R04 needs to take "risk assessment" as the core output dimension of the summary results, with a scenario adaptability of ≥92%. For work meetings, D02 needs to focus on extracting "action items, time nodes, and responsible parties," and R04 needs to highlight the description of action items to ensure that the summary results meet the usage needs of the scenario.

[0093] like Figure 2 As shown, a lightweight optimization device for AI models suitable for resource-constrained devices includes: The hardware resource analysis module 201 is used to extract device resource constraint threshold data for the computing power, memory, and power consumption hardware parameters of the embedded device RK3588, complete the construction of the hardware resource adaptation feature set, and provide hardware adaptation basis for subsequent model optimization. The meeting data processing module 202 is used to collect raw voice and text data from the meeting scenario, and sequentially perform data cleaning, format standardization, and feature extraction preprocessing operations to generate a meeting data sample set suitable for AI model training and inference. The model lightweight optimization module 203 is used to perform customized lightweight optimization on the initial AI meeting summary model based on the hardware resource adaptation feature set. By pruning redundant network layers, quantizing model parameters, and compressing model size, a lightweight AI meeting summary model adapted to the characteristics of RK3588 hardware resources is generated. The edge model inference module 204 is used to deploy the lightweight and optimized AI meeting summary model to the RK3588 device, import the pre-processed meeting data sample set for model inference calculation, and generate preliminary meeting summary results. The summary result verification module 205 is used to verify the meeting summary results output by the model, and to correct and optimize the results by combining them with the actual meeting content, so as to generate automatic meeting summary data that meets the actual needs. The process integration solution module 206 is used to integrate the entire process of lightweight model optimization with the conference data processing inference link, complete the collaborative adaptation of each link, and form an intelligent processing solution for end-side conference data that is compatible with RK3588.

[0094] A computing device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute any AI model lightweight optimization method applicable to resource-constrained devices.

[0095] The methods and / or embodiments in this application can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by a processing unit, it performs the functions defined in the methods of this application.

[0096] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0097] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application.

Claims

1. A lightweight optimization method for AI models suitable for resource-constrained devices, characterized in that, include: For the computing power, memory, and power consumption hardware parameters of the embedded device RK3588, extract device resource constraint threshold data and construct a hardware resource adaptation feature set; Collect raw voice and text data from meeting scenarios, perform data cleaning, format standardization, and feature extraction preprocessing to generate a meeting data sample set suitable for AI model training and inference; Based on the hardware resource adaptation feature set, the initial AI meeting summary model is customized and optimized to be lightweight. Redundant network layers are removed, model parameters are quantized, and model size is compressed to adapt to the characteristics of RK3588 hardware resources, thus generating a lightweight AI meeting summary model adapted to the characteristics of RK3588 hardware resources. The lightweight and optimized AI meeting summary model was deployed to the RK3588 device, and the pre-processed meeting data sample set was imported for model inference to generate meeting summary results. The model outputs meeting summary results for data validation, and the results are corrected and optimized based on the actual meeting content to generate automatic meeting summary data that meets actual needs. By integrating the lightweight optimization process of the model with the inference link of the meeting data processing, an intelligent processing solution for end-side meeting data adapted to RK3588 is formed.

2. The method according to claim 1, characterized in that, Raw audio and text data from meeting scenarios are collected, cleaned, standardized in format, and preprocessed with feature extraction to generate a meeting data sample set suitable for AI model training and inference, including: To address the resource constraints of computing power, memory, and power consumption of the embedded device RK3588, and in conjunction with the low-redundancy data requirements for lightweight training and inference of edge AI meeting summary models, real-time audio stream data and multi-format raw text data in meeting scenarios were collected in a targeted manner to form multi-source heterogeneous raw data materials for meetings with scene tags. The collected heterogeneous raw data materials from multiple sources of the meeting were subjected to edge-adaptive preprocessing operations, which included redundant and noisy data cleaning, multi-format data standardization and normalization, and lightweight semantic feature extraction. This process eliminated invalid and redundant information, unified edge-side model to adapt to data format, and mined core semantic features that fit the needs of meeting summary. Based on the lightweight adaptation standard of edge AI model corresponding to the hardware resource characteristics of RK3588, the resource adaptability compliance verification of the cleaned, standardized and feature-extracted meeting data is carried out to screen out the effective meeting data that meets the requirements of low computing power training and low memory inference of edge model. The validated meeting data is integrated with lightweight features and a dedicated dataset is constructed for the edge. The data is divided into layers and packaged in a lightweight manner according to the edge resource allocation requirements for model training and inference, generating a meeting data sample set suitable for AI model training and inference.

3. The method according to claim 1, characterized in that, Based on the hardware resource adaptation feature set, the initial AI meeting summary model is customized and lightweightly optimized by pruning redundant network layers, quantizing model parameters, and compressing the model size to adapt to the characteristics of RK3588 hardware resources, including: The hardware resource analysis algorithm is used to deeply analyze the computing power capacity threshold, memory usage limit, and power consumption control standard of RK3588, and generate the hardware adaptation baseline rules and parameter optimization threshold range for AI models deployed on the edge. By connecting the network structure data and hardware resource adaptation feature set of the model summarized from the initial AI conference, a lightweight optimization benchmark for the model is constructed. The lightweight optimization rules for model layers, parameters, and volume are determined through model-hardware matching logic. At the same time, a customized lightweight adaptation model is introduced. Combined with the model compression algorithm and the RK3588 hardware resource profile, a dynamic computation mechanism of demand-constraint-optimization adaptation is established. Using model optimization as the coverage dimension, the optimization range and hardware resource adaptation data corresponding to network layers, model parameters, and model volume are integrated, combined with model hardware adaptation baseline rules and parameter optimization thresholds. By integrating model structure and hardware constraint data through lightweight processing mechanisms such as network layer pruning, parameter quantization, and volume compression, and combining customized lightweight adaptation models to enhance the accuracy of optimization and adaptation, a lightweight AI conference summary model adapted to the hardware resource characteristics of RK3588 is generated by combining the dual-layer optimization architecture characteristics of simplified model structure and efficient parameter compression.

4. The method according to claim 3, characterized in that, The lightweight, optimized AI meeting summary model was deployed to the RK3588 device. Preprocessed meeting data samples were imported for model inference, generating meeting summary results, including: Based on the deployment goals of RK3588, the performance requirements of AI model inference, and the adaptability of meeting summary scenarios, we integrate the model deployment adaptation parameters, sample data import specifications, inference resource allocation thresholds, and summary result output dimension information to clarify the setting scope and collaboration requirements of each operation item. Based on the collaborative requirements of equipment resource utilization efficiency and the effectiveness of meeting summary reasoning results, the deployment reasoning logic is designed, and the execution standards of core operation items and auxiliary operation items are clarified. Among them, the core operation items include the edge deployment and debugging of the lightweight AI meeting summary model and the targeted import of pre-processed meeting data samples into reasoning. The auxiliary operation items include dynamic scheduling of computing power during the reasoning process and real-time matching of semantic features of meeting data. In response to the consistency requirements of summary results in various lightweight meeting processing scenarios, inference optimization rules are set up to dynamically adapt deployment parameters, verify compliance of sample import, adjust resource allocation in real time, and accurately adapt feature matching. The system integrates the requirements for setting operation items, coordinating resource utilization and inference result effectiveness, and inference optimization rules to generate a standardized edge-side meeting inference base dataset containing deployment type, execution specifications, collaborative logic, and optimization strategies. It reads the model deployment parameters, sample data types, and inference resource usage information of the dataset, sorts them according to the correlation between data and inference accuracy and meeting summary efficiency, completes the model edge-side inference operation, and generates meeting summary results.

5. The method according to claim 1, characterized in that, By integrating the lightweight model optimization process with the conference data processing inference chain, an intelligent end-side conference data processing solution adapted to RK3588 is formed, including: The entire process of AI-powered meeting summary is categorized and broken down, including hardware resource analysis, meeting data preprocessing, model lightweight optimization, edge deployment and inference, and summary result verification and optimization. This process generates process link division results, a link interaction association table, and data flow mapping relationships. Based on the RK3588 computing power, memory, and power consumption constraints and the actual needs of lightweight meeting processing, the process link division results, link interaction association table, and data flow mapping relationship are standardized. The process integration granularity is set in combination with the device resource adaptation characteristics, and the link linkage response cycle is clarified according to the real-time requirements of meeting summary. Based on the process integration granularity standard and the link linkage response cycle, each process link is grouped and integrated to form a core system for end-side conference processing that includes hardware resource adaptation sub-process, conference data processing sub-link, lightweight model construction sub-process, and end-side inference optimization sub-link. Prioritize the execution of each core system for conferencing on each endpoint. Combine the dynamic resource allocation rules of RK3588 with the performance requirements of AI model endpoint operation, formulate a full-process collaborative operation plan, and generate an intelligent conferencing data processing solution for endpoints adapted to RK3588 that includes system number, process composition details, execution timing plan, resource adaptation rules, and scenario adaptation requirements.

6. A lightweight optimization device for AI models suitable for resource-constrained devices, characterized in that, The device is used to execute executable instructions to perform the lightweight optimization method for AI models applicable to resource-constrained devices as described in any one of claims 1 to 5.

7. An electronic device, characterized in that, include: First processor; And a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the AI ​​model lightweight optimization method for resource-constrained devices according to any one of claims 1 to 5 by executing the executable instructions.

8. A computing device, the device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, When the computer program instructions are executed by the processor, the device is triggered to execute the AI ​​model lightweight optimization method applicable to resource-constrained devices as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Civil aircraft four-cabin anomaly detection method under resource limited condition

    CN118395827A

  • Audio quality analysis method and system based on algorithm

    CN121171209A

  • Multi-computing-power collaborative video structured analysis method

    CN121415318A

  • AI digital human conference proxy method and device under off-line local area network and medium

    CN121585789A