This invention discloses a lightweight AI underlying optimization method and
system that significantly saves national computing resources, involving the fields of AI underlying
algorithm optimization, domestic computing power scheduling, and embedded AI chips. Based on proportional filtering logic and STE dynamic lossless quantization encoding, this invention achieves end-to-end underlying reconstruction of AI models, incorporating domestically produced lightweight chips and
inference models, and setting thresholds for redundancy removal, computing power utilization, accuracy loss, and computing power reduction. Through feature filtering, operator reconstruction, weight quantization,
chip operator alignment, and accuracy self-calibration,
model compression and efficient computing power scheduling are achieved, eliminating reliance on foreign technologies. This
system can reduce AI computing
power consumption in key national sectors by more than 65%, achieve a computing power
utilization rate of ≥85%, increase
inference speed by 2 times, and reduce accuracy loss by ≤0.5%. It is applicable to government, industry, cloud, and
edge computing scenarios, contributing to the intensive, green, and autonomous use of computing power, and filling the gap in domestic AI underlying lightweight optimization technology.