Vertical knowledge model training method and device based on AI training and pushing integrated machine
By constructing a multimodal training dataset using an AI training and promotion integrated machine and performing hybrid precision quantization and fine-grained feature enhancement, the problems of data privacy and training efficiency in existing technologies are solved, enabling efficient training and sharing of high-precision bird identification and professional knowledge output.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN CHUANGYINGXIN IND CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies cannot efficiently train high-precision bird vertical knowledge models on desktop devices while ensuring the privacy and security of users' original data, and they cannot meet the requirements for fine-grained recognition accuracy, professional knowledge matching, and shooting scene adaptability.
A multimodal training dataset is built using an AI training and push integrated machine. The model is fine-tuned using a hybrid precision quantization and fine-grained feature enhancement mechanism. Closed-loop optimization is performed by combining user feedback data, enabling efficient local training and open-source sharing of the model.
While protecting user data privacy, it improves the fine-grained accuracy of bird identification and the ability to output professional knowledge, meeting the needs of efficient, low-cost model training and open-source sharing.
Smart Images

Figure CN122113991A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence hardware equipment technology, and in particular relates to a vertical knowledge model training method and device based on an AI training and promotion integrated machine. Background Technology
[0002] Currently, with the rapid development of AI technology, in-depth research and application of vertical knowledge models in specific interest areas have become a new hot topic.
[0003] Taking bird research as an example, the group of bird observation and photography enthusiasts continues to expand. Related communities have accumulated a large amount of unstructured data, such as original bird photographs, identification experience, living environment and habit records. Users have a continuous demand for high-precision, fine-grained bird identification and professional vertical knowledge services.
[0004] For bird photography enthusiasts, based on their own original bird images and unstructured text and image data from interest communities, it is impossible to complete the training of local high-precision bird vertical knowledge models at low cost and high efficiency through desktop devices while ensuring the privacy and security of private original shooting data. Furthermore, the existing general-purpose large models cannot meet the professional use and open-source sharing needs of enthusiasts in terms of fine-grained species recognition accuracy, vertical professional knowledge matching degree, and shooting scene adaptability in the bird sub-field. Summary of the Invention
[0005] This application provides a vertical knowledge model training method and device based on an AI training and promotion integrated machine. By constructing a complete process from multi-source data processing, model lightweighting, fine-tuning training, professional evaluation, open source conversion to feedback-driven closed-loop optimization, it can efficiently train a high-precision bird vertical knowledge model locally on a desktop device while ensuring the privacy and security of users' original data. This improves fine-grained recognition accuracy and scene adaptability, meeting the needs of professional use and open source sharing.
[0006] In a first aspect, embodiments of this application provide a method for training a vertical knowledge model based on an AI training and induction integrated machine, comprising: acquiring multi-source unstructured data of a target research object, the target research object including birds; processing the multi-source unstructured data of the target research object to construct a multimodal training dataset of the target research object; performing model lightweighting processing on a bird knowledge base model based on the multimodal training dataset; using the lightweighted bird knowledge base model, performing model fine-tuning training based on a fine-grained feature enhancement mechanism to obtain a bird knowledge multimodal model; performing model evaluation on the bird knowledge multimodal model according to professional dimensions related to bird identification, and iteratively optimizing the bird knowledge multimodal model based on the evaluation results to obtain a qualified model; receiving user inference requests and feedback data based on the open-source model derived from the qualified model, completing closed-loop incremental optimization, and obtaining an optimized bird knowledge professional model.
[0007] Secondly, this application provides a vertical knowledge model training device based on an AI training and inference integrated machine, comprising: an acquisition unit for acquiring multi-source unstructured data of a target research object, the target research object including birds; a processing unit for processing the multi-source unstructured data of the target research object to construct a multimodal training dataset of the target research object; performing model lightweighting processing on a bird knowledge base model based on the multimodal training dataset; using the lightweighted bird knowledge base model, performing model fine-tuning training based on a fine-grained feature enhancement mechanism to obtain a bird knowledge multimodal model; performing model evaluation on the bird knowledge multimodal model according to professional dimensions related to bird identification, and iteratively optimizing the bird knowledge multimodal model based on the evaluation results to obtain a qualified model; receiving user inference requests and feedback data based on the open-source model derived from the qualified model, completing closed-loop incremental optimization, and obtaining an optimized bird knowledge professional model.
[0008] Thirdly, embodiments of this application provide a server including a processor, a memory, a communication interface, and a computer program, the computer program being stored in the memory and configured to be executed by the processor, the program including instructions for performing steps in the method as described in any one of the first aspects.
[0009] As can be seen in this embodiment, a multimodal training dataset is constructed by processing multi-source unstructured data from the community using an AI training and induction integrated machine. Based on this dataset, a bird knowledge base model is lightweighted, and then fine-tuned through a fine-grained feature enhancement mechanism. After professional evaluation and iteration, a qualified model is obtained and converted to open-source format for standardized export. Simultaneously, user inference requests and feedback data are used to achieve closed-loop incremental optimization of the model. This solution can efficiently complete local model training on desktop devices while protecting the privacy of users' original data, reducing resource consumption, improving the model's fine-grained recognition accuracy and adaptability to professional scenarios, and enabling continuous iterative upgrades of model capabilities to meet the needs of professional training and open-source sharing in the field of bird recognition. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 Two schematic diagrams showing the appearance and structure of the AI training and promotion integrated machine provided in the embodiments of this application; Figure 2 A flowchart illustrating a vertical knowledge model training method based on an AI training and promotion integrated machine, provided in an embodiment of this application; Figure 3 A structural block diagram of another AI training and promotion integrated machine provided in the embodiments of this application; Figure 4 A flowchart illustrating another vertical knowledge model training method based on an AI training and promotion integrated machine provided in this application embodiment; Figure 5 This is a schematic diagram illustrating the multimodal dataset construction process provided in the embodiments of this application; Figure 6 This is a schematic diagram of the model lightweighting process provided in the embodiments of this application; Figure 7 This is a schematic diagram of the model fine-tuning training process provided in the embodiments of this application; Figure 8 This is a schematic diagram of the model evaluation and iterative optimization process provided in the embodiments of this application; Figure 9 A functional unit structure diagram of a vertical knowledge model training device based on an AI training and promotion integrated machine provided in this application embodiment; Figure 10 This is a schematic diagram of the server structure provided in an embodiment of this application. Detailed Implementation
[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0013] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but in some embodiments includes steps or units not listed, or in some embodiments includes other steps or units inherent to these processes, methods, products, or apparatuses.
[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0015] In the embodiments of this application, "and / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone; A and B exist simultaneously; B exists alone. Among them, A and B can be singular or plural.
[0016] In this embodiment, the symbol " / " can indicate that the preceding and following objects are in an "or" relationship. Alternatively, the symbol " / " can also represent a division sign, i.e., performing a division operation. For example, A / B can mean A divided by B.
[0017] In the embodiments of this application, "at least one item" or its similar expression refers to any combination of these items, including any combination of a single item or a plurality of items. "One or more" means one or more, while "multiple" means two or more. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.
[0018] In the embodiments of this application, "equal to" can be used with "greater than" and is applicable to technical solutions used when "greater than" is used; it can also be used with "less than" and is applicable to technical solutions used when "less than" is used. When "equal to" is used with "greater than", it is not used with "less than"; when "equal to" is used with "less than", it is not used with "greater than".
[0019] The number of birdwatching and photography enthusiasts continues to grow, and the community has accumulated a wealth of original bird photographs, species identification experience, habitat and habit records, and other first-hand data, as well as unstructured text and image data from the community. At the same time, there is a strong demand for high-precision fine-grained bird identification and professional vertical knowledge output.
[0020] Existing bird identification model training and application technologies suffer from numerous insurmountable shortcomings: general-purpose large models can only perform basic identification of common birds, with extremely low accuracy in identifying closely related, easily confused bird species and regionally endemic subspecies. They also cannot output specialized vertical knowledge such as bird identification points, shooting techniques, and habitat habits. Professional bird identification models are mostly trained by research institutions based on fixed open-source datasets, failing to integrate first-hand shooting data and practical identification experience from hobbyist communities, resulting in extremely poor adaptability to actual shooting scenarios. In the model training stage, cloud training requires uploading user-generated shooting data and private community data, posing irreversible risks of data leakage and copyright infringement. Consumer-grade PCs lack sufficient computing power to support full-scale fine-tuning of multimodal bird vertical models, leading to long training cycles and a tendency for GPU memory overflow. Professional training and push servers are costly, bulky, and have high maintenance barriers, making them unsuitable for desktop deployment and low-cost usage requirements.
[0021] Furthermore, while existing desktop AI training and promotion all-in-one machines have the capability to deploy large models, they lack a dedicated training solution for bird enthusiast communities. This makes it impossible to achieve automated processing of unstructured data from the community, efficient local training of fine-grained bird models, standardized export of open-source models, and closed-loop optimization based on community feedback. Consequently, they cannot meet the core needs of bird enthusiast communities to train high-precision bird vertical models at low cost and high efficiency, and to achieve open-source sharing, while ensuring data privacy and security.
[0022] To address the aforementioned technical challenges, this application implements comprehensive technical improvements based on a desktop AI training and promotion all-in-one machine. This device enables end-to-end local model training, eliminating the need to upload private text and image data from the community to the cloud, thus fundamentally protecting data privacy and copyright security. By automating the processing of unstructured, multi-source data from bird communities, a multimodal training dataset for birds, adapted to real-world shooting scenarios, is constructed, resolving the issues of fragmented and low-quality vertical training data. High-speed cross-system data synchronization and integrity verification between the Windows interaction layer and the Linux computing layer are achieved, supporting hot-swapping without restarts and without performance loss, meeting the high-efficiency operation requirements of desktop hardware. A hybrid precision quantization combined with automatic optimization strategies for computing power and storage is adopted for the bird knowledge base model, adapting to the hardware specifications of the training and promotion all-in-one machine, effectively avoiding memory overflow and significantly shortening the training cycle. Fine-grained feature enhancement and contrastive learning mechanisms are used for model fine-tuning, significantly improving the model's accuracy in identifying easily confused bird species and its ability to output professional vertical knowledge. Simultaneously, the model's open-source format is standardized for export, and a closed-loop incremental optimization system is built based on user feedback, enabling the model to continuously iterate and upgrade based on community data.
[0023] To further clarify and fully illustrate the technical solution of this application, the following detailed description of various embodiments of the vertical knowledge model training method and device based on the AI training and promotion integrated machine is provided in conjunction with specific examples.
[0024] The following is combined with Figure 1 The structure of the AI training and promotion all-in-one machine shown is as follows: Figure 1 These are schematic diagrams illustrating two different appearances and structures of the AI training and promotion all-in-one machine provided in embodiments of this application. This application uses a desktop-level AI training and promotion all-in-one machine as its core hardware carrier, and its appearance is as follows... Figure 1 As shown in Appearance 1 and Appearance 2, its hardware architecture adopts a layered design, consisting of a computing core layer, a multi-system interaction and collaboration layer, a high-speed I / O and storage expansion layer, and a stable power supply and environment adaptation layer. The hardware characteristics of each layer are deeply integrated with the technical solution of this application, providing full-process hardware support for the training, optimization, deployment and closed-loop iteration of the bird vertical knowledge model.
[0025] The core computing layer provides the necessary computing power for model training and inference, meeting the computing and storage requirements for local deployment of large models, mixed-precision quantization, and fine-grained feature enhancement and fine-tuning. In one possible example, this integrated training and inference machine is equipped with an AI processor with a computing power of 126 TOPS, natively supporting the local deployment of large models with 70B / 120B parameters, and is configured with up to 128GB of LPDDR5 high-capacity memory and up to 4TB×3 NVMe high-speed SSD storage. This hardware configuration provides sufficient computing and storage resources for the rapid loading of bird multimodal datasets, mixed-precision quantization of models (INT4 quantization for the general semantic layer and INT8 quantization for the fine-grained feature layer), and fine-grained feature enhancement and fine-tuning, effectively solving the technical pain points of insufficient computing power and memory overflow in desktop devices, and ensuring the smooth implementation of lightweight processing and efficient training of bird knowledge base models.
[0026] The multi-system interaction and collaboration layer is used to realize cross-system data synchronization and process communication, ensuring efficient collaboration between interactive operations and computing power execution in the model training process. This training and push-integrated machine is natively compatible with both Windows and Linux systems, constructing the Windows interaction layer and Linux computing power layer in this application. At the hardware level, it supports cross-system data transmission based on a shared memory mechanism, combined with data integrity verification using hash algorithms, breakpoint resumption, and cross-system process communication methods, achieving high-speed and secure synchronization of the bird multimodal training dataset between the Windows interaction layer and the Linux computing power layer. This provides stable support for processing multi-source unstructured data from the community and for cross-system execution of model training tasks.
[0027] The high-speed I / O and storage expansion layer is used to meet the needs of multi-source data import, model file reading and writing, and peripheral device interaction, covering the entire data interaction process of model training. In a possible example, this training and push machine is equipped with a Type-C 3.2 port, a gigabit Ethernet port, and a multi-channel USB 3.2 Type-A high-speed interface, while also supporting three M.2 NVMe SSD expansion interfaces. The above hardware configuration can realize the rapid import of multi-source data such as original images taken by bird enthusiasts, unstructured chat records from social media groups, and shared text and image materials. It can also support the hierarchical storage of training sets, validation sets, and test sets, as well as high-speed reading and writing and standardized export of model checkpoints and open-source format model files, adapting to the efficient operation of bird multimodal dataset construction, model evaluation, and export processes.
[0028] A stable power supply and environment adaptation layer is used to provide a stable operating environment for long-term model training and adapt to abnormal handling scenarios during training. In one possible example, this training and push machine adopts a 19V DC power supply scheme, supports a wide power supply range of 110-240V, and has the ability to adapt to a wide temperature range of 0-40℃, as well as an intelligent temperature control and heat dissipation system. During the model fine-tuning training phase, the hardware can avoid performance fluctuations caused by abnormal high temperatures through temperature control. Combined with stable power supply and peripheral connection protection, it can achieve stable execution of breakpoint resume training, training task scheduling, and computing resource allocation, providing reliable hardware environment support for the entire model training process.
[0029] It should be noted that the above hardware configuration is only a specific example provided in this application and does not limit the hardware structure of the AI training and promotion all-in-one machine. The AI training and promotion all-in-one machine may also include other equivalent hardware modules and configurations that can achieve the same function.
[0030] In summary, the technical solution of this application is deeply integrated with the hardware architecture of the aforementioned AI training and promotion all-in-one machine. It uses the upper-level layered hardware architecture as the overall framework and the lower-level specific hardware configuration as the implementation means. It accurately matches the hardware capabilities such as computing power support, cross-system collaboration, data interaction, and stability assurance with the entire process of vertical knowledge model training. Ultimately, it realizes local privacy-preserving training, efficient optimization, standardized export, and closed-loop iteration of community feedback for bird vertical multimodal models on desktop hardware, effectively solving many defects in existing technologies.
[0031] Based on the hardware architecture of the aforementioned desktop AI training and promotion all-in-one machine, the specific implementation steps of the vertical knowledge model training method based on the AI training and promotion all-in-one machine described in this application will be explained in detail below.
[0032] Please see Figure 2 , Figure 2 This application provides a flowchart illustrating a vertical knowledge model training method based on an AI training and promotion integrated machine, which is applied to, for example... Figure 1 The AI training and promotion integrated machine shown includes the following steps S201-S206: Step S201: Obtain multi-source unstructured data of the target research object's community, including birds.
[0033] Among them, the community refers to online communities for bird observation, photography, species identification, and popular science exchange; multi-source unstructured data refers to bird-related data from such communities that do not have a unified format or organization, including but not limited to bird image data taken by bird enthusiasts, unstructured chat records in the community, bird image and text materials shared within the community, and text describing bird species and habits.
[0034] Step S202: Process the multi-source unstructured data of the community to construct a multimodal training dataset of the target research object.
[0035] This step is used to standardize the above multi-source unstructured data to form a multimodal training dataset suitable for training bird vertical knowledge models.
[0036] In specific implementations, in some embodiments, the process of processing the multi-source unstructured data of the community to construct a multimodal training dataset for the target research object includes: importing, parsing, and cleaning the multi-source unstructured data of the community to obtain a valid image dataset and a valid text corpus. The multi-source unstructured data of the community includes bird image data originally taken by bird enthusiasts, unstructured chat log files of the community, and bird image and text materials shared within the community; performing multimodal image-text semantic alignment, fine-grained annotation, hierarchical partitioning of the dataset, and wide-area expansion of the dataset on the valid image dataset and the valid text corpus to obtain a multimodal training dataset.
[0037] In some embodiments, the process of importing, parsing, and cleaning the multi-source unstructured data from the community to obtain a valid image dataset and a valid text corpus includes: importing and parsing multi-format image and text data; performing structured extraction on the EXIF information of the image data; performing word segmentation and named entity recognition on the chat text; extracting bird-related valid text based on a bird-specific vocabulary; performing deduplication on the images using a perceptual hashing algorithm; removing invalid images using a fuzzy detection algorithm; performing deduplication on the text using a text similarity algorithm; and persistently storing the processed valid data.
[0038] In practice, the AI training and promotion all-in-one machine receives all the raw data uploaded by users in batches through the community data import and management panel and completes the format parsing. Among them, image data formats include but are not limited to JPG, PNG, RAW, TIFF, and BMP, and text data formats include but are not limited to TXT, JSON, CSV, DOCX, and PDF. The format compatibility parsing module automatically extracts image files and text content.
[0039] The AI training and promotion integrated machine performs structured extraction of EXIF information from image data. EXIF information refers to image metadata recorded in the Exchangeable Image File Format, which is additional information automatically written to the image file by the capturing device. The extraction process is represented as follows: ; Where I is the original bird image, EXIF(·) is the EXIF information extraction function, P is the extracted set of shooting parameters, f is the shooting focal length, F is the aperture value, T is the shutter speed, L is the shooting location, and D is the shooting time. This converts the unstructured EXIF metadata into a standardized key-value pair format, forming a basic image attribute ledger.
[0040] Meanwhile, the AI training and push integrated machine performs word segmentation and named entity recognition (NER) on the chat text. Its processing logic is as follows: It uses a word segmentation algorithm to break down continuous text into independent lexical units. Combined with a pre-set bird-specific vocabulary database (containing bird species names, morphological characteristic words, habitat terms, identification key words, etc.), it filters and extracts valid bird-related text through word matching and entity annotation, filtering out irrelevant chatter, advertising information, garbled characters, meaningless statements, and other invalid content to ensure the vertical relevance of the text data. The calculation formula is: ; in, V represents the original chat text content, V represents the built-in bird-specific vocabulary database, NER(·) is the named entity recognition function, and T represents the extracted bird-related valid text. The AI training and promotion all-in-one machine thus eliminates irrelevant content such as invalid chat, advertisements, and garbled characters, ensuring the vertical relevance of the text data.
[0041] Subsequently, the AI training and promotion integrated machine uses a perceptual hashing algorithm to perform deduplication on the images. It calculates the Hamming distance between the image hash values to determine similarity and removes duplicate images with similarity exceeding a set threshold. Simultaneously, it employs a fuzzy detection algorithm based on a formula... Image blurriness is calculated, and images with blurriness below a threshold are deemed invalid and discarded. A text similarity algorithm is then employed, using the cosine similarity formula. Calculate the similarity between texts and remove duplicate and meaningless text content. Here, V is the image blur index; a higher value indicates a clearer image, and a lower value indicates a blurrier image. is the Laplacian operator (second-order differential operator) used to detect edge and detail changes in an image; Var(·) is the variance function, which calculates the variance of all pixel values in the matrix. For vectors and The cosine similarity, with values ranging from [ [1,1], the closer to 1, the more similar they are.
[0042] Finally, the AI training and promotion integrated machine stores the processed valid image data and valid text data into its high-speed storage module for persistent storage, resulting in a standardized and unified valid image dataset and valid text corpus, providing high-quality data support for subsequent multimodal semantic alignment, fine-grained annotation, and model training.
[0043] In some embodiments, the multimodal image-text semantic alignment, fine-grained annotation, hierarchical partitioning of the dataset, and wide-area expansion of the dataset for the effective image dataset and the effective text corpus include: mapping the effective images in the effective image dataset and the effective text in the effective text corpus to a unified semantic space and performing semantic pairing and alignment; performing automatic pre-annotation on the semantically aligned image-text samples, and then correcting the pre-annotation results through expert verification to obtain a gold standard annotated dataset; dividing the gold standard annotated dataset into a training set, a validation set, and a test set according to the stratified sampling principle, with the species distribution and easily confused species ratio of each subset being consistent with the gold standard annotated dataset; and fusing the training set with an open-source bird dataset and performing data augmentation on small sample species to obtain the bird multimodal training dataset.
[0044] In practice, firstly, multimodal image-text semantic alignment is performed on valid images and valid text. The goal is to establish a one-to-one correspondence between visual information in images and semantic information in text, enabling the model to learn the association patterns between bird image features and text descriptions, providing consistent supervision signals for subsequent training. Specifically, features are extracted from both images and text and mapped to the same semantic space. Cosine similarity is used to calculate the degree of image-text matching, and the image with the highest similarity that meets the threshold is selected to form a pair with the text. The calculation formula is as follows: ; in, , These are the image feature vector and the text feature vector, respectively, and S is the image-text similarity. When S is greater than or equal to a preset threshold, it is determined to be a semantic match, and an aligned image-text sample is formed.
[0045] Subsequently, automatic pre-annotation was performed, uniformly labeling the aligned image and text samples. The annotations included: bird species labels, easily confused feature regions, and key identification parts in the images, as well as species names, morphological features, habitat information, and identification points in the text. After automatic pre-annotation, community bird experts verified, corrected, and supplemented the annotation results, removing erroneous annotations and improving missing annotations, ultimately forming the gold standard annotation dataset. This dataset contains: semantically aligned image-text pairing samples, accurate species labels, fine-grained feature region annotations, standardized text entity annotations, and expert verification labels, and can be directly used for model training and evaluation.
[0046] Next, the gold standard labeled dataset was divided using stratified sampling to ensure that the species distribution and common mixing percentages of the training, validation, and test sets were consistent with the overall dataset. The division relationship is expressed as follows: ; Where D is the gold standard labeled dataset, and Split(·) is the hierarchical partitioning process. , , The sets are, in order: training set, validation set, and test set.
[0047] Finally, the training set was merged with publicly available bird datasets to supplement the species coverage; data augmentation was performed on species with small sample sizes, such as image flipping, brightness adjustment, and text synonym replacement, to expand sample diversity, improve model generalization ability, and ultimately form a standardized multimodal bird training dataset.
[0048] Based on the specific embodiments of step S201 above, please refer to... Figure 5 , Figure 5 This is a schematic diagram of the multimodal dataset construction process provided in the embodiments of this application, such as... Figure 5 As shown, this process starts with user / community uploaded data to build a standardized multimodal training dataset for birds. The specific process is as follows: First, users / communities upload image data and text data respectively, resulting in the original image set and the original text corpus. The original image set enters the image data cleaning stage, where it undergoes perceptual hashing for deduplication and fuzzy detection for removal, yielding deduplicated images and invalid images respectively. These two sets are then merged into a valid image dataset. The original text corpus enters the text data cleaning and filtering stage, where it undergoes text similarity deduplication and NER+ bird vocabulary filtering, yielding deduplicated text and invalid text respectively. These two sets are then merged into a valid text corpus.
[0049] Next, the valid image dataset is structured by extracting EXIF information to generate an image attribute ledger; the valid text corpus is annotated with semantic entities to generate a standardized text corpus. The image attribute ledger and the standardized text corpus then enter the multimodal image-text semantic alignment stage, and after cosine similarity matching, an aligned image-text sample set is obtained.
[0050] Subsequently, the aligned image and text sample set was automatically pre-annotated to generate annotation results, which were then verified and corrected by experts to form the gold standard annotated dataset. Next, through stratified sampling, the dataset was divided into training, validation, and test sets. Finally, open-source data fusion and few-shot augmentation were performed on the training set to obtain the final multimodal bird training dataset.
[0051] This application adopts a Windows+Linux dual-system collaborative architecture. The Windows interactive layer is responsible for providing a visual operation interface and undertaking front-end interactive functions such as data access, manual verification, annotation correction, and dataset construction and management. The Linux computing power layer is responsible for model loading, computing power scheduling, model training and inference acceleration. The two layers work together through shared memory and high-speed data channels.
[0052] In each embodiment of step S202 above, all operations described in step S202, such as data import, parsing, cleaning, multimodal image-text semantic alignment, fine-grained annotation, dataset partitioning and wide-area expansion, are performed in the Windows interactive layer. The multimodal training dataset obtained after processing is stored in the storage area corresponding to the Windows interactive layer.
[0053] To fully utilize the dedicated computing resources of the Linux computing layer and ensure model training efficiency and hardware stability, it is necessary to securely and completely synchronize the multimodal training dataset built in the Windows interactive layer to the Linux computing layer, and to perform integrity verification on the transmitted data to ensure that the data is not lost, damaged, or tampered with during the cross-system migration process, thus providing a reliable data foundation for subsequent model training.
[0054] Based on this, this application sets step S202 to process the multi-source unstructured data of the community and construct the multimodal training dataset of the target research object. Then, it sets a step to perform cross-system data synchronization and integrity verification of the multimodal training dataset. Specifically, this includes: transmitting the multimodal training dataset of the Windows interaction layer of the AI training and push all-in-one machine to the Linux computing power layer of the AI training and push all-in-one machine based on a shared memory mechanism.
[0055] This application utilizes a high-speed cross-system data synchronization module with a dual-system collaborative kernel layer, based on a shared memory mechanism, to rapidly synchronize the full bird training dataset stored in the Windows interactive layer to a dedicated training directory in the Linux computing layer. The synchronization transfer rate is no less than 10GB / s. The synchronization process is represented as follows: ; in, The full bird training dataset is stored in the Windows interactive layer. This refers to the physical address of the shared memory in the Linux layer, and Copy(·) is the high-speed copy function for shared memory. This is to synchronize the full training dataset to the Linux layer.
[0056] In some embodiments, a hash algorithm is used to perform data integrity verification on the files during data transmission; if the verification result fails, the files that failed the verification are resumed from the point of interruption until all files pass the verification; the data synchronization status and integrity verification result are synchronized to the Windows interaction layer through cross-system process communication.
[0057] During data synchronization, this application performs integrity verification on each file based on the MD5 hash algorithm, checking whether the number of files, file size, hash value, and the source file in the Windows layer are consistent. The verification rules are as follows: ; Among them, MD5(·) is the MD5 hash value calculation function, and Check(·) is the data integrity verification function, which outputs the result of successful or failed verification.
[0058] If the verification result fails, this application immediately triggers the breakpoint resume mechanism, only retransmitting the files that failed the verification, without retransmitting the full data, until the full file verification passes; after synchronization is completed, a data verification report is generated in the Linux layer, and at the same time, the synchronization status and verification report are synchronized to the Windows interactive layer in real time through the cross-system process communication module, updating the data status and triggering the subsequent training process.
[0059] Through the above steps, this application finally outputs the full bird training dataset, data integrity verification report, and data synchronization completion status signal, which are synchronized and verified by the Linux computing layer, providing a reliable and complete data foundation for subsequent model training.
[0060] Step S203: Perform model lightweighting processing on the bird knowledge base model based on the multimodal training dataset.
[0061] The lightweighting process includes performing mixed-precision quantization on the model and optimizing the allocation of storage and computing resources required for model operation.
[0062] The verified dataset is a multimodal bird training dataset that is securely synchronized from the Windows interaction layer to the Linux computing layer via high-speed transmission through shared memory of the dual systems and MD5 integrity verification. This dataset is completely consistent with the source data built in the Windows interaction layer and serves as the direct data basis for subsequent model loading and lightweight processing.
[0063] In some embodiments, the lightweighting process of the bird knowledge base model based on the multimodal training dataset includes: automatically matching and loading the bird knowledge base model according to the species information and sample type information of the multimodal training dataset; selecting a mixed precision quantization strategy according to the hardware configuration information of the AI training and push integrated machine, performing INT4 quantization on the general semantic layer of the bird knowledge base model, and performing INT8 quantization on the bird fine-grained feature layer of the bird knowledge base model; allocating contiguous memory and GPU memory space according to the volume of the quantized model, and performing unified model loading and GPU memory scheduling optimization.
[0064] The multimodal training dataset includes species information and sample type information. Species information includes bird species names, number of species, distribution of easily confused closely related species, and sample size distribution. Sample type information includes image resolution, shooting scene, text knowledge type, annotation granularity, and data format.
[0065] This application automatically matches and loads bird knowledge base models from the bird knowledge base model library based on the species information and sample type information of the verified dataset. The matching process is based on the species coverage and sample complexity of the dataset. The matching principle is to prioritize multimodal basic models that have learned a large number of bird image and text features in the pre-training stage. At the same time, the model version with corresponding parameter amount is matched according to the scale information such as the total number of samples and species in the dataset, so that the semantic ability of the model is adapted to the species features and text description scenarios of the dataset.
[0066] For example, if the dataset contains only 30 common forest birds and the total number of samples is 2,000, the AI training and promotion machine will automatically match a lightweight bird knowledge base model with 7B parameters; if the dataset covers 200 bird species, includes a large number of easily confused birds of prey and waterfowl and the number of samples exceeds 10,000, the AI training and promotion machine will automatically match a bird knowledge base model with 13B parameters.
[0067] This application selects a hybrid precision quantization strategy based on the hardware configuration information of the AI training and inference integrated machine. The core of hybrid precision quantization lies in: by differentially quantizing the weight parameters of different levels of the bird knowledge base model, it can compress the model size and optimize hardware resource consumption while retaining the parameter precision required for fine-grained recognition. The bird knowledge base model includes a general semantic layer responsible for basic language and visual feature processing, and a fine-grained bird feature layer responsible for subtle feature recognition. The model weight parameters are the core parameters used by the model to perform weighted calculations on the input bird image and text features, extract fine-grained features such as plumage color, markings, beak shape, and tail shape, and complete species classification. Hybrid precision quantization is the quantization processing of these weight parameters. The quantized weight parameters will directly replace the original full-precision weights for subsequent model training and inference.
[0068] Specifically, the general semantic layer of the bird knowledge base model is responsible for basic language and visual feature processing, and can use lower precision quantization to compress the size. The fine-grained bird feature layer is responsible for recognizing subtle differences such as plumage color, markings, beak shape, and tail shape, and needs to retain higher precision to ensure recognition results. For example, when the AI training and inference integrated machine is equipped with 96GB of video memory and a high-performance GPU, it automatically selects the INT4+INT8 layered quantization strategy. Here, INT4 refers to 4-bit quantization and INT8 refers to 8-bit quantization. The lower the number of quantization bits, the smaller the model size and the higher the inference efficiency, but the feature expression precision is correspondingly reduced. INT4 quantization is performed on the general semantic layer to significantly reduce the model's footprint, while INT8 quantization is performed on the bird fine-grained feature layer to ensure the recognition accuracy of easily confused bird species.
[0069] The AI training and promotion all-in-one machine automatically selects an INT4 / INT8 mixed-precision quantization strategy based on the overall hardware configuration to quantize and compress the basic model, optimizing memory usage while preserving the parameter accuracy of the bird fine-grained feature extraction layer, ensuring the model's recognition capability. This quantization process is achieved through the following formula: ; in, is the full-precision (FP32) weight parameter of the original model; b is the quantization bit width, the general semantic layer uses INT4 quantization (b=4), and the bird fine-grained feature layer uses INT8 quantization (b=8); Min() and Max() are the minimum and maximum value functions of the weight parameters; Round() is the rounding function. These are the quantized model weight parameters. After quantization, Will replace Participate in the subsequent training and inference process of the model to achieve lightweight deployment of the model.
[0070] This application is based on the volume of the quantized model (determined by the quantized weight parameters). The overall memory footprint determines the allocation of contiguous memory and GPU memory, and performs unified model loading and GPU memory scheduling optimization to avoid GPU memory overflow and memory fragmentation during training, ensuring stable and efficient model operation. For example, if the quantized model's GPU memory usage is approximately 10GB, the AI training and push all-in-one machine allocates 12GB of contiguous space in the total GPU memory and reserves 2GB as a batch training cache. High-frequency model weights are placed in high-speed GPU memory, while low-frequency data is stored in main memory, achieving hierarchical scheduling and efficient loading.
[0071] Based on the various embodiments of step S203 described above, please refer to... Figure 6 , Figure 6 This is a schematic diagram of the model lightweighting process provided in an embodiment of this application. Figure 6As shown, step S203 performs lightweight processing on the bird knowledge base model based on the validated bird multimodal training dataset. The specific process is as follows: First, using the validated multimodal bird training dataset as input, the basic model is automatically matched and loaded. By matching species information / sample information, a suitable bird knowledge base model is obtained.
[0072] Subsequently, the mixed-precision quantization strategy selection stage is entered: INT4 quantization is performed on the general semantic layer of the model, and INT8 quantization is performed on the fine-grained feature layer of the model. The results of the two quantization processes are used to generate the quantized model.
[0073] Next, continuous memory / GPU memory space is allocated to the quantized model, followed by unified model loading and GPU memory scheduling optimization, and finally the lightweight model is output.
[0074] As can be seen, in this embodiment, by automatically matching the bird knowledge base model based on the dataset information, the model and training task can be accurately adapted; by using INT4 and INT8 hierarchical mixed precision quantization, the model size can be significantly compressed and the computing power consumption can be reduced while effectively preserving the fine-grained recognition accuracy of birds; by allocating continuous memory according to the model size and optimizing memory scheduling, memory overflow and memory fragmentation can be avoided, ensuring stable and efficient loading and operation of the model. Overall, the lightweight processing is automated, efficient and reliable.
[0075] After the aforementioned lightweighting process, although the model has been significantly optimized in terms of size, memory usage, and runtime efficiency, the lightweight bird knowledge base model still only possesses general visual and linguistic feature capabilities. It lacks sufficient learning of fine-grained differences in bird species such as plumage coloration, markings, beak shape, and tail shape, and its adaptability to domain-specific multimodal knowledge is weak, making it difficult to directly meet the needs of high-precision bird identification and understanding. Therefore, targeted deep fine-tuning and feature enhancement are still needed to further improve the model's accuracy and generalization ability in bird-specific scenarios.
[0076] Therefore, this application continues with the model fine-tuning training steps: Step S204: Using the lightweight bird knowledge base model, fine-tuning training is performed based on a fine-grained feature enhancement mechanism to obtain a multimodal bird knowledge model.
[0077] Among them, fine-grained features include key features that are subtle but distinguishable between bird species, such as plumage distribution, pattern shape, beak morphology, tail feather outline, eye markings, wing patch structure, foot features, and body proportions.
[0078] Fine-grained feature enhancement mechanisms are used to improve the model's ability to capture and express the aforementioned subtle distinguishing features, enabling the model to focus on key differences between easily confused species and improve the accuracy of image-text multimodal matching and species identification.
[0079] In some embodiments, the fine-grained feature enhancement mechanism-based model fine-tuning training includes: configuring a corresponding training strategy according to the bird fine-grained recognition scenario, wherein the training strategy includes training parameters and data augmentation strategies, and the bird fine-grained recognition scenario includes a closely related species differentiation scenario, an individual recognition scenario of different ages or sexes, a recognition scenario under different shooting conditions, and a multimodal knowledge matching scenario; constructing a total loss function including classification loss and contrastive loss, wherein the classification loss is used for basic bird species identification and classification, and the contrastive loss is used to enhance the feature discrimination of samples of different categories; and iteratively updating the model parameters according to the training strategy and the total loss function using contrastive learning and hard example optimization methods until a preset convergence condition is met, thereby completing the model fine-tuning training.
[0080] This application addresses the identification characteristics of bird species that are highly similar and easily confused. It configures corresponding training strategies according to the fine-grained bird identification scenario. By comparing and learning, it brings the feature distance between samples of the same type of bird closer together and widens the feature distance between samples of different species. At the same time, by mining difficult examples, it focuses on strengthening the model's learning ability for closely related bird species that are similar in morphology and difficult to distinguish, thereby further improving the accuracy of fine-grained identification.
[0081] The fine-grained bird recognition scenarios specifically include scenarios for distinguishing closely related species, identifying individuals of different ages or sexes, recognizing individuals under different shooting conditions, and multimodal knowledge matching. For example, the scenario for distinguishing closely related species involves differentiating between sparrows and tree sparrows, and egrets and intermediate egrets, which are very similar in appearance; the scenario for recognizing individuals of different ages or sexes involves distinguishing between adult birds and juvenile birds, and between male and female birds based on their plumage color differences; the scenario for recognizing individuals under different shooting conditions involves recognizing bird images taken in low light, backlight, occlusion, and at long distances; and the scenario for multimodal knowledge matching involves associating and matching bird images with multimodal knowledge such as morphological descriptions, habitual texts, and distribution areas.
[0082] To address the different scenarios mentioned above, this application configures the training strategy by adjusting the learning rate, batch size, loss function weights, and data augmentation strategies. For example, for complex scenarios such as distinguishing closely related species, the AI training and push integration machine automatically adopts a smaller initial learning rate and increases the weight of the contrastive learning loss to enhance the model's ability to capture subtle differences. For low-light recognition scenarios, data augmentation strategies such as random brightness adjustment and noise simulation are enabled to improve the model's robustness to complex shooting conditions.
[0083] Building upon this, this application employs a model training mechanism combining contrastive learning and hard example mining. Contrastive learning uses a loss function to shorten the feature vector distance between samples of the same class and widen the feature vector distance between samples of different classes, enabling the model to simultaneously learn species classification and fine-grained feature differentiation. Its loss function is: ; Where L represents the total loss, which is used to guide the overall model training; This represents the classification loss, used to ensure the model outputs the correct species category; α represents the contrast loss, used to enhance the feature distinction between similar species; α represents a fixed weighting coefficient, used to balance the two losses.
[0084] The specific formula for calculating classification loss is as follows: ; Where n represents the number of samples used in one training session; m represents the number of bird species categories; This indicates whether the i-th sample belongs to the j-th class; 1 indicates yes, 0 indicates no. This represents the probability that the model determines the i-th sample belongs to the j-th class. The purpose of this loss is to enable the model to identify the bird species corresponding to each image as accurately as possible.
[0085] The specific formula for calculating the contrast loss is as follows: ; Where β is the set boundary threshold; This represents the feature similarity between the i-th sample and other samples of the same class. This represents the feature similarity between the i-th sample and samples from different classes. The purpose of this loss is to make the features of birds of the same class more similar, and to make the features of different birds, especially closely related species that are prone to mixing, more distant.
[0086] Under the constraints of the aforementioned loss function, contrastive learning uses the loss function to narrow the feature vector distance between samples of the same class and widen the feature vector distance between samples of different classes, enabling the model to learn more discriminative feature representations. For example, when training sparrow samples, the model is guided to narrow the feature distance with other sparrow samples while widening the feature distance with tree sparrow samples that have a similar appearance, thereby strengthening the feature boundaries between the two.
[0087] Difficult example mining dynamically filters difficult examples such as those that the model misidentifies or has low confidence, such as those that are prone to confusion, low-quality images, and rare morphologies. It increases their weight in training, allowing the model to focus on learning error-prone points. For example, when the model frequently misidentifies a blackbird as a crow, the AI training and induction machine will automatically label these samples as difficult examples, increase their sampling frequency and loss weight, and force the model to strengthen its learning of key distinguishing features such as the blackbird's yellow beak and eye rings.
[0088] Under the constraints of the aforementioned loss function, contrastive learning guides the model to optimize the feature distances between similar and dissimilar samples, thereby iteratively updating the model parameters. Difficult example mining specifically strengthens the model's learning of error-prone samples. The two work synergistically to achieve continuous optimization of model parameters and complete model fine-tuning training. The iteratively updated model parameters mainly include the weight parameters at each level of the model, specifically the weights of the general semantic layer and the bird fine-grained feature layer. These are of the same type as the weight parameters in the mixed precision quantization mentioned earlier and can directly adapt to the model training requirements. The iterative update process involves continuously optimizing these weight parameters, enabling the model to gradually acquire high-precision, fine-grained bird recognition capabilities.
[0089] The preset convergence conditions are: within a preset number of iterations, the total loss function value tends to stabilize and the change is less than a set threshold, and / or the model's recognition accuracy on the validation set tends to stabilize and the improvement is less than a set threshold, and / or the model iteration reaches the preset maximum number of training iterations.
[0090] In some embodiments, after performing model fine-tuning training based on the fine-grained feature enhancement mechanism, the method further includes: performing training task scheduling and computing resource allocation; freezing the general semantic layer and fine-tuning the bird fine-grained feature layer; monitoring training process data and hardware operating status in real time; performing hardware temperature control when abnormal hardware temperature is detected; performing breakpoint continuation training when training is abnormally interrupted; saving model checkpoints according to iteration rounds; and selecting the model with the best validation index as the bird knowledge multimodal model.
[0091] During training, training task scheduling and computing resource allocation are performed. This involves rationally allocating computing resources based on model training requirements and hardware capacity to ensure efficient and stable training while avoiding resource waste or overload. Freezing the model's general semantic layer means fixing the already trained general feature extraction capabilities of the underlying model (i.e., retaining the basic visual and language-related general knowledge already mastered by the model without changing its parameters), preventing subsequent training from disrupting existing general capabilities. Fine-tuning the bird-specific feature layer means updating only the parameters related to bird-specific features in this layer for bird recognition scenarios. These parameters include weight parameters for extracting fine-grained morphological features such as plumage color, markings, beak shape, tail shape, and body shape, as well as classification and mapping parameters for bird species differentiation and multimodal knowledge matching. This allows the model to focus on learning bird-specific features and associating with multimodal knowledge, enhancing its adaptability to bird recognition scenarios while retaining general capabilities. Simultaneously, it improves training efficiency, reduces unnecessary computing power consumption, and ensures the model can accurately capture bird-specific features while efficiently completing training tasks.
[0092] To ensure stable and reliable training, this application monitors training data such as loss value and accuracy, as well as hardware operating status such as GPU temperature, load, and memory usage in real time. When excessively high hardware temperature is detected, it automatically reduces power consumption and performs hardware temperature control measures such as frequency reduction and current limiting to prevent hardware damage. When training is unexpectedly interrupted, it automatically resumes training from the breakpoint based on the most recently saved model state, eliminating the need to start training from scratch and improving training continuity and efficiency. During training, this application saves model checkpoints for each iteration and evaluates the model for each iteration based on metrics such as recognition accuracy and recall on the validation set. The model with the best validation metrics is selected as the final multimodal bird knowledge model, ensuring that the model has the best generalization ability and practicality in fine-grained bird recognition and multimodal understanding tasks.
[0093] Based on the above embodiments of step S204, please refer to Figure 7 , Figure 7 This is a schematic diagram of the model fine-tuning training process provided in an embodiment of this application. Figure 7 As shown, this step is based on the lightweight bird knowledge base model, and the model depth is fine-tuned through a fine-grained feature enhancement mechanism to obtain a multimodal bird knowledge model.
[0094] First, based on fine-grained bird recognition scenarios such as distinguishing closely related species, identifying different ages / sexes, complex shooting conditions, and multimodal knowledge matching, corresponding training strategies are configured to generate training strategies that include learning rate, batch size, loss function weights, and data augmentation strategies. On this basis, a training mechanism combining contrastive learning and hard example mining is constructed. Classification loss L1 and contrastive loss L2 are calculated separately, and backpropagation is performed based on the total loss L to update the model parameters.
[0095] Subsequently, training task scheduling and computing resource allocation are executed. During training, the underlying general semantic layer is frozen, and only the fine-grained feature layer for birds is fine-tuned before entering the model iterative training phase. During iterative training, the training process and hardware operating status are monitored in real time. If abnormal hardware temperature is detected, temperature control and frequency reduction are implemented. If training is interrupted, breakpoint resumption is performed.
[0096] Meanwhile, during the training process, model checkpoints are saved according to the iteration rounds, and the models of each round are evaluated based on the validation set to select the model with the best validation index, and finally output the bird knowledge multimodal model.
[0097] As can be seen, in this embodiment, considering the high similarity and easy confusion among bird species, a training strategy is configured to adapt to the fine-grained recognition scenario. Contrastive learning and hard example mining are combined to enhance the model's ability to learn subtle features. By rationally allocating computing resources, freezing the underlying general semantic layer, and fine-tuning the bird-specific feature layer, the model's basic capabilities are preserved while focusing on learning bird-specific features, thus improving training efficiency and recognition accuracy. Real-time monitoring of hardware status, temperature control protection, and breakpoint continuation training ensure a stable, safe, and continuous training process. By saving model checkpoints iteratively and selecting the optimal model, the final bird knowledge multimodal model is ensured to have reliable, stable, and optimal performance in bird fine-grained recognition and multimodal understanding tasks.
[0098] After obtaining the multimodal model of bird knowledge through the above fine-tuning training, in order to further verify the actual effect of the model in the task of fine-grained bird recognition and multimodal understanding, and to ensure that the model's accuracy, stability, and practicality meet the requirements of practical applications, it is necessary to conduct professional evaluation and iterative optimization. Therefore, this application continues with the following steps: Step S205: Perform model evaluation on the bird knowledge multimodal model according to professional dimensions related to bird identification, and iteratively optimize the bird knowledge multimodal model based on the evaluation results to obtain a qualified model.
[0099] In some embodiments, the step of evaluating the bird knowledge multimodal model according to professional dimensions related to bird identification and iteratively optimizing the model based on the evaluation results includes: performing inference tests on the bird knowledge multimodal model based on the test set to generate multi-dimensional vertical category capability evaluation indicators; comparing the multi-dimensional vertical category capability evaluation indicators with preset thresholds to determine whether the bird knowledge multimodal model meets the standards; if it does not meet the standards, performing targeted optimization processing according to the defect type of the non-compliant indicators, and then re-performing the step "performing inference tests on the bird knowledge multimodal model based on the test set" and subsequent steps on the bird knowledge multimodal model; if it meets the standards, performing version archiving processing on the compliant model, and synchronizing the model evaluation results and compliance status across systems to the Windows interaction layer.
[0100] In this embodiment, the AI training and inference integrated machine takes the optimal model weight file output in step S205 and the test set data output in step S202 as input, performs inference tests on the bird knowledge multimodal model, and generates evaluation indicators covering the four core dimensions of bird vertical ability.
[0101] The first dimension is the fine-grained species identification accuracy, which includes the overall identification accuracy and the accuracy for distinguishing easily confused species. The overall identification accuracy A is calculated using the following formula: ; Where TP is the number of true positives, representing the number of samples correctly identified as target birds by the model; TN is the number of true negatives, representing the number of samples correctly identified as non-target birds by the model; FP is the number of false positives, representing the number of non-target samples that the model misclassifies as target birds; and FN is the number of false negatives, representing the number of target samples that the model misses as non-target birds. The accuracy rate for identifying easily confused closely related bird species... Calculate using the following formula: ; in, To identify the correct number of samples in easily mixed populations. This represents the number of identification errors in easily mixed samples.
[0102] The second dimension is the vertical category knowledge matching degree, and the knowledge matching accuracy K is calculated using the following formula: ; Where C represents the number of bird habits, identification points, and habitat information output by the model that meet expert standards, and T represents the total number of vertical knowledge contents output by the model.
[0103] The third dimension is scene adaptability, which is the recognition accuracy of the statistical model on real-world images with different lighting, focal lengths, and angles, and is statistically analyzed separately for each shooting scene; the fourth dimension is inference performance, which includes the inference latency and memory usage of a single image.
[0104] After generating the above evaluation indicators, the AI training and inference integrated machine compares them with preset compliance thresholds. The preset compliance thresholds are: overall species identification accuracy ≥ 95%, easily confused species identification accuracy ≥ 90%, vertical category knowledge matching degree ≥ 92%, and single image inference latency ≤ 100ms. If all indicators are met, the model is deemed qualified, a complete multi-indicator model evaluation report for bird-specific dimensions is generated, the qualified model is archived, and the evaluation report and model compliance status are synchronized to the model management panel in the Windows interactive layer through the cross-system process communication module.
[0105] If any metric fails to meet the standard, the model is deemed unqualified. The AI training and inference system automatically analyzes the type of metric defect and matches corresponding iterative optimization strategies: if the accuracy of easily confused species identification is insufficient, it automatically supplements comparison samples of easily confused species, strengthens the contrastive learning strategy, increases the number of training iterations, and initiates incremental training; if the vertical knowledge matching degree is insufficient, it supplements the corresponding species' professional text corpus, optimizes the image-text alignment strategy, and fine-tunes the model's semantic layer parameters; if the inference latency is insufficient, it re-executes quantization optimization, adopts a lower bit-width quantization strategy, and prunes redundant parameters of the model. After targeted optimization, the AI training and inference system re-executes the inference testing and evaluation process on the model until all metrics of the model meet the standards, finally obtaining the final version of the bird knowledge multimodal model with satisfactory metrics.
[0106] Based on the various embodiments of step S205 described above, please refer to... Figure 8 , Figure 8 This is a schematic diagram illustrating the model evaluation and iterative optimization process provided in an embodiment of this application. Figure 8 As shown, this step evaluates and iteratively optimizes the multimodal model of bird knowledge according to the professional dimensions related to bird identification, and finally obtains a qualified model.
[0107] First, the multimodal bird knowledge model is tested using a test set to generate four evaluation metrics: fine-grained species identification accuracy, vertical category knowledge matching degree, shooting scene adaptability, and inference performance. These four evaluation metrics are then compared with preset thresholds to determine if the model meets all the criteria.
[0108] If the model is deemed unqualified, targeted optimizations are performed based on the type of defect: for insufficient accuracy in identifying easily confused types, supplementary comparison samples are added and contrastive learning is strengthened; for insufficient vertical knowledge matching, professional text corpora are added and image-text alignment is optimized; for insufficient inference performance, quantization optimization is re-executed and the model is pruned. After optimization, the process returns to the step of performing inference tests based on the test set, and evaluation and judgment are performed again until the model meets all qualification conditions.
[0109] If all models are deemed to meet the standards, version archiving is performed on the compliant models, and the model evaluation results and compliance status are synchronized across systems to the Windows interaction layer. Finally, the final compliant model that meets the preset standards is output.
[0110] As can be seen, in this embodiment, by conducting multi-index evaluation and targeted iterative optimization according to the professional dimensions of bird identification, on the one hand, it can comprehensively and accurately measure the model's real performance in fine-grained recognition, vertical knowledge matching, scene adaptability, and inference performance, avoiding evaluation bias caused by relying on a single index; on the other hand, the automated optimization strategy for different defect types can quickly locate and make up for the model's shortcomings, ensuring that the model meets the preset application standards in terms of recognition accuracy, professional knowledge matching degree, robustness to complex scenes, and inference efficiency, providing a reliable guarantee for the subsequent actual deployment and stable operation of the model.
[0111] Based on the above evaluation and optimization, the model has reached the preset performance standards. In order to facilitate rapid deployment and use on different platforms and in different application scenarios, this application continues to perform the following steps: converting the qualified model into open source format and exporting it in a standardized manner.
[0112] In some embodiments, the process of converting and exporting the compliant model to an open-source format includes: converting the compliant model to multiple mainstream open-source common formats; automatically generating model metadata and open-source documentation, and associating the open-source documentation with the converted open-source format model file; synchronizing the open-source format model file and its accompanying model metadata and open-source documentation across systems to the Windows interactive layer; and performing local inference deployment verification on the open-source model synchronized to the Windows interactive layer to confirm that the capabilities of the open-source model are consistent with those of the compliant model.
[0113] In this embodiment, the AI training and inference integrated machine automatically converts the qualified model weight files in PyTorch format into two mainstream open-source universal formats, GGUF and ONNX, using a standardized conversion function. This ensures compatibility with mainstream open-source inference frameworks such as Ollama and LMStudio, allowing the public to load and use them without any barriers. The conversion process is achieved through the following formula: ; in, For the standardization model, S is the target open-source format set, and F(·) is the model format standardization conversion function. , This is the converted open-source format model file.
[0114] Subsequently, the AI training and inference all-in-one machine automatically generates open-source metadata and documentation for the model, including model version, species coverage, core capability indicators, usage methods, inference environment requirements, and training dataset source descriptions. It also links these metadata and documentation with the converted model files, making it easy for users to quickly understand the model's capabilities and usage methods.
[0115] Next, the AI training and push integrated machine synchronizes the open-source format model files and accompanying documents from the Linux computing power layer to the Windows interaction layer through a high-speed cross-system data synchronization module. At the same time, it encrypts and archives the original model weight files to ensure the security of the original model version.
[0116] Finally, the AI training and inference integrated machine automatically completes the local inference service deployment and verification of the open-source format model, tests the model's recognition accuracy, inference latency, and knowledge output capability, and ensures that the open-source model and the trained qualified model are completely consistent in capability with no loss of accuracy, providing a reliable foundation for the subsequent public use of the model and closed-loop optimization based on community feedback.
[0117] Step S206: Receive user reasoning requests and feedback data based on the open-source model derived from the qualified model, complete closed-loop incremental optimization, and obtain the optimized bird knowledge professional model.
[0118] In some embodiments, the step of receiving user inference requests and feedback data based on the open-source model derived from the benchmark model to complete closed-loop incremental optimization includes: receiving a bird identification inference request initiated by a user; processing the bird images in the bird identification inference request using the open-source model to output species identification results and bird vertical knowledge information; collecting user feedback correction data on the identification results; performing expert verification on the collected feedback correction data; filtering valid feedback samples and storing them in a feedback sample library; counting the number of valid feedback samples in the feedback sample library in real time; determining whether to trigger incremental training based on the statistical results; when incremental training is triggered, processing the valid feedback samples in the feedback sample library; and re-executing the step "using the lightweight processed bird knowledge base model, performing model fine-tuning training based on a fine-grained feature enhancement mechanism to obtain a bird knowledge multimodal model" and subsequent steps based on the current open-source model to complete the version update of the open-source model and obtain the optimized bird knowledge professional model.
[0119] In this embodiment, the AI training and inference integrated machine first receives bird identification inference requests initiated by the public and community users through an open-source inference framework, i.e., bird images uploaded by users. After the model automatically performs standardized preprocessing on the images, it outputs vertical knowledge content containing species identification results, confidence scores, identification points, habit information, and shooting suggestions. The inference latency for a single image does not exceed 100ms. The inference process can be represented as follows: ; in, For inference requests based on bird images uploaded by users, M is the open-source model, L is the bird species name predicted by the model, I is the species identification, habits, and shooting-related vertical knowledge output by the model, and C is the confidence level of the prediction result, with a value ranging from 0 to 1.
[0120] Subsequently, the AI training and promotion integrated machine collects feedback and correction data from community users and the public on the recognition results, including the original requested image, model prediction results, user-corrected real species labels, and supplementary vertical knowledge information; then, community bird experts verify the feedback data, remove erroneous feedback, retain valid feedback samples and store them in the feedback sample library, forming a continuously updated incremental training dataset.
[0121] The AI training and promotion all-in-one machine counts the number of valid samples in the feedback sample library in real time. When the number of valid samples accumulates to a preset threshold (e.g., 50 samples), or when the number of easily confused misidentified samples accumulates to a preset threshold (e.g., 20 samples), the incremental training process is automatically triggered, and an incremental training trigger signal is generated. If the threshold is not reached, feedback data is continuously collected.
[0122] When the AI training and push integrated machine receives the incremental training trigger signal, it automatically performs data augmentation and dataset partitioning on the valid samples in the feedback sample library according to the process in step S201, and synchronizes them to the Linux computing power layer. Based on the current open source model, it sequentially executes the entire process of fine-tuning training, evaluation, and open source format conversion from step S204 to step S206 to complete the automatic version update of the model. The update process does not affect the normal use of the existing open source model, thereby realizing the continuous iterative optimization of the model's capabilities.
[0123] As can be seen, in this embodiment, by receiving user inference requests, collecting and verifying feedback data, and triggering incremental training according to preset conditions, closed-loop optimization of the model is achieved. This not only utilizes real-world feedback to continuously strengthen the model's shortcomings, but also ensures the quality of incremental data through expert verification. Ultimately, without affecting existing services, the model's performance is continuously and stably improved through iterative improvements.
[0124] Please see Figure 3 , Figure 3 This is a structural block diagram of another AI training and promotion all-in-one machine provided in an embodiment of this application. The AI training and promotion all-in-one machine described in this application is divided into three layers according to its functional architecture: Windows interaction layer, Linux computing power layer, and hardware module. Each layer works together to provide support for the whole process training of bird vertical knowledge model.
[0125] Among them, the Windows interaction layer, as the user-facing front-end interaction and management layer, is mainly used to provide a visual operation interface, complete the access of multi-source data in the community, data parsing and cleaning, image and text semantic alignment, fine-grained annotation and dataset construction, support expert verification, parameter configuration, task control, and display data synchronization status, model training process, evaluation results and model achievement information. It is also responsible for open source model management, receiving user inference requests and feedback data collection, realizing full-process human-computer interaction and data management.
[0126] The Linux computing layer serves as the core of background computing execution and model processing. It is mainly responsible for efficient computing tasks related to the model, including loading the bird knowledge base model, mixed precision quantization and lightweight processing, performing model fine-tuning training, computing resource scheduling, comparative learning and hard example mining, conducting model inference testing, professional dimension evaluation and iterative optimization, completing the open source format conversion of the model and version archiving, and performing incremental training based on user feedback samples to achieve closed-loop iterative updates of the model, ensuring the efficient and stable operation of model training, inference and optimization.
[0127] The hardware module serves as the physical foundation and operational guarantee of the entire machine, providing underlying hardware support for dual-system collaboration, data transmission, and model computation. It provides computing power support through AI processors, high-speed video memory, and large-capacity memory, enables multi-source data import and fast reading and writing of model files through high-speed storage and I / O interfaces, supports high-speed data interaction and process communication between Windows and Linux dual systems through a shared memory mechanism, and provides a safe, reliable, and continuous operating environment for long-term model training through stable power supply, intelligent temperature control, and wide temperature adaptability.
[0128] Please see Figure 4 , Figure 4 The flowchart illustrates another method for training a vertical knowledge model based on an AI training and promotion integrated machine, as provided in this application embodiment. The method includes the following steps S401-S410: Step S401, Upload Personal Dataset in Vertical Domain: Users upload multi-source unstructured data on bird communities to the Windows interactive layer of the all-in-one machine as the raw data source for model training.
[0129] Step S402, Data Import and Preprocessing: The Windows interactive layer completes data import, format parsing, and data cleaning to obtain a valid image dataset and a valid text corpus.
[0130] Step S403, Visual Data Labeling and Sample Optimization: The Windows interactive layer completes multimodal image-text semantic alignment, automatic pre-labeling, expert verification and correction, divides the training set, validation set and test set by stratified sampling, and integrates open source data and performs small sample data augmentation to construct a standardized bird multimodal training dataset.
[0131] Step S404, Cross-system data synchronization and verification: The Windows interaction layer synchronizes the completed multimodal training dataset to the Linux computing layer at high speed through a shared memory mechanism, and performs integrity verification based on a hash algorithm. If the verification fails, the interruption resumes automatically to ensure that the data is complete and tamper-free.
[0132] Step S405, Basic Model Loading and Lightweight Optimization: The Linux computing layer matches and loads the bird knowledge basic model based on the dataset information, calls the computing power and storage resources of the hardware module, performs INT4 / INT8 mixed precision quantization and memory scheduling optimization, and completes the model lightweighting process.
[0133] Step S406, Small Sample Model Fine-Tuning Training and Monitoring: The Linux computing power layer performs model fine-tuning training based on the fine-grained feature enhancement mechanism. It enhances the model's ability to identify easily confused bird species through comparative learning and difficult sample mining. At the same time, it freezes the underlying general semantic layer and only fine-tunes the bird fine-grained feature layer. It monitors the training process and hardware status in real time, and performs temperature control protection and breakpoint resume training until training is completed.
[0134] Step S407, Model Evaluation and Iterative Optimization: The Linux computing power layer conducts multi-dimensional professional evaluation based on the test set, generating fine-grained recognition accuracy, vertical knowledge matching degree, shooting scene adaptability and inference performance indicators; compare the indicators with preset thresholds, and if they do not meet the standards, perform targeted optimization according to the defect type, retrain and evaluate, until all indicators of the model meet the standards.
[0135] Step S408, Model Export and Local Inference Deployment: The Linux computing layer converts the qualified model into mainstream open-source formats such as GGUF and ONNX, generates supporting metadata and documentation, synchronizes it to the Windows interactive layer, and completes local inference deployment verification to ensure that the open-source model has the same capabilities as the original qualified model.
[0136] Step S409, User initiates inference request / feedback: The user calls the inference service through the Windows interaction layer to initiate a bird recognition request, and can submit feedback to correct the recognition results.
[0137] Step S410, Inference Service Call and Incremental Training Trigger: The Windows interaction layer receives the user's inference request, calls the open-source model to output recognition results and vertical category knowledge; at the same time, it collects user feedback data, stores it in the feedback sample library after expert verification; when the effective sample size or the number of easily confused error samples accumulates to a preset threshold, incremental training is automatically triggered, and steps S406 to S410 are re-executed based on the feedback samples to achieve closed-loop iterative optimization of the model's capabilities.
[0138] As can be seen in this embodiment, a multimodal training dataset is constructed by processing multi-source unstructured data from the community using an AI training and induction integrated machine. Based on this dataset, a bird knowledge base model is lightweighted, and then fine-tuned through a fine-grained feature enhancement mechanism. After professional evaluation and iteration, a qualified model is obtained and converted to open-source format for standardized export. Simultaneously, user inference requests and feedback data are used to achieve closed-loop incremental optimization of the model. This solution can efficiently complete local model training on desktop devices while protecting the privacy of users' original data, reducing resource consumption, improving the model's fine-grained recognition accuracy and adaptability to professional scenarios, and enabling continuous iterative upgrades of model capabilities to meet the needs of professional training and open-source sharing in the field of bird recognition.
[0139] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the server includes the corresponding hardware structure and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0140] This application embodiment can divide the server into functional units according to the above method example. For example, each function can be divided into different functional units, or two or more functions can be integrated into one processing module. The integrated unit can be implemented in hardware or as a software program module. It should be noted that the unit division in this application embodiment is illustrative and only represents a logical functional division, while other division methods may be used in actual implementation.
[0141] In the case of using integrated units, please refer to Figure 9 , Figure 9 A functional unit structure diagram of a vertical knowledge model training device based on an AI training and promotion integrated machine provided in this application embodiment is shown below. Figure 9 As shown, the vertical knowledge model training device 9 based on the AI training and promotion integrated machine includes: Acquisition unit 901 is used to acquire multi-source unstructured social data of the target research object, including birds; Processing unit 902 is used to process the multi-source unstructured data of the community to construct a multimodal training dataset of the target research object; perform model lightweighting on the bird knowledge base model based on the multimodal training dataset; use the lightweighted bird knowledge base model to perform model fine-tuning training based on a fine-grained feature enhancement mechanism to obtain a bird knowledge multimodal model; perform model evaluation on the bird knowledge multimodal model according to professional dimensions related to bird identification, and perform iterative optimization on the bird knowledge multimodal model based on the evaluation results to obtain a qualified model; receive user inference requests and feedback data based on the open-source model derived from the qualified model, complete closed-loop incremental optimization, and obtain an optimized bird knowledge professional model.
[0142] As can be seen in this embodiment, a multimodal training dataset is constructed by processing multi-source unstructured data from the community using an AI training and induction integrated machine. Based on this dataset, a bird knowledge base model is lightweighted, and then fine-tuned through a fine-grained feature enhancement mechanism. After professional evaluation and iteration, a qualified model is obtained and converted to open-source format for standardized export. Simultaneously, user inference requests and feedback data are used to achieve closed-loop incremental optimization of the model. This solution can efficiently complete local model training on desktop devices while protecting the privacy of users' original data, reducing resource consumption, improving the model's fine-grained recognition accuracy and adaptability to professional scenarios, and enabling continuous iterative upgrades of model capabilities to meet the needs of professional training and open-source sharing in the field of bird recognition.
[0143] In some embodiments, in processing the multi-source unstructured data of the community to construct a multimodal training dataset for the target research object, the processing unit 902 is configured to: import, parse, and clean the multi-source unstructured data of the community to obtain an effective image dataset and an effective text corpus. The multi-source unstructured data of the community includes bird image data originally taken by bird enthusiasts, unstructured chat log files of the community, and bird image and text materials shared within the community; perform multimodal image-text semantic alignment, fine-grained annotation, hierarchical partitioning of the dataset, and wide-area expansion of the dataset on the effective image dataset and the effective text corpus to obtain a multimodal training dataset.
[0144] In some embodiments, in terms of performing multimodal image-text semantic alignment, fine-grained annotation, hierarchical partitioning of the dataset, and wide-area expansion of the dataset on the effective image dataset and the effective text corpus, the processing unit 902 is configured to: map the effective images in the effective image dataset and the effective text in the effective text corpus to a unified semantic space and perform semantic pairing and alignment; perform automatic pre-annotation on the image-text samples after semantic alignment, and then correct the pre-annotation results through expert verification to obtain the gold standard annotated dataset; divide the gold standard annotated dataset into a training set, a validation set, and a test set according to the stratified sampling principle, with the species distribution and easily confused species ratio of each subset being consistent with the gold standard annotated dataset; perform fusion of the training set with the open-source bird dataset, and perform data augmentation on small sample species to obtain the bird multimodal training dataset.
[0145] In some embodiments, after processing the multi-source unstructured data of the community to construct the multimodal training dataset of the target research object, the processing unit 902 is further configured to: transmit the multimodal training dataset of the Windows interactive layer of the AI training and push all-in-one machine to the Linux computing power layer of the AI training and push all-in-one machine based on a shared memory mechanism.
[0146] In some embodiments, in the process of lightweighting the bird knowledge base model based on the multimodal training dataset, the processing unit 902 is configured to: automatically match and load the bird knowledge base model based on the species information and sample type information of the multimodal training dataset; select a mixed precision quantization strategy based on the hardware configuration information of the AI training and push integrated machine, perform INT4 quantization on the general semantic layer of the bird knowledge base model, and perform INT8 quantization on the bird fine-grained feature layer of the bird knowledge base model; allocate contiguous memory and video memory space according to the volume of the quantized model, and perform unified model loading and video memory scheduling optimization.
[0147] In some embodiments, in performing model fine-tuning training based on a fine-grained feature enhancement mechanism, the processing unit 902 is configured to: configure a corresponding training strategy according to the bird fine-grained recognition scenario, the training strategy including training parameters and data augmentation strategy, the bird fine-grained recognition scenario including a closely related species differentiation scenario, a recognition scenario of individuals of different ages or sexes, a recognition scenario under different shooting conditions, and a multimodal knowledge matching scenario; construct a total loss function including classification loss and contrast loss, wherein the classification loss is used for basic bird species identification and classification, and the contrast loss is used to enhance the feature discrimination of samples of different categories; and iteratively update the model parameters according to the training strategy and the total loss function using contrastive learning and hard example optimization methods until a preset convergence condition is met, thereby completing the model fine-tuning training.
[0148] In some embodiments, in evaluating the bird knowledge multimodal model according to bird identification-related professional dimensions and iteratively optimizing the bird knowledge multimodal model based on the evaluation results, the processing unit 902 is configured to: perform inference testing on the bird knowledge multimodal model based on the test set to generate multi-dimensional vertical category capability evaluation indicators; compare the multi-dimensional vertical category capability evaluation indicators with preset thresholds to determine whether the bird knowledge multimodal model meets the standards; if it does not meet the standards, perform targeted optimization processing according to the defect type of the non-compliant indicators, and then re-execute the step "perform inference testing on the bird knowledge multimodal model based on the test set" and subsequent steps; if it meets the standards, perform version archiving processing on the compliant model and synchronize the model evaluation results and compliance status across systems to the Windows interaction layer.
[0149] In some embodiments, in receiving user inference requests and feedback data based on the open-source model derived from the benchmark model, completing closed-loop incremental optimization, and obtaining an optimized bird knowledge professional model, the processing unit 902 is configured to: receive a bird identification inference request initiated by a user; process the bird image in the bird identification inference request using the open-source model; output species identification results and bird vertical category knowledge information; collect user feedback correction data on the identification results; perform expert verification on the collected feedback correction data; filter valid feedback samples and store them in the feedback sample library; count the number of valid feedback samples in the feedback sample library in real time; determine whether to trigger incremental training based on the statistical results; when incremental training is triggered, process the valid feedback samples in the feedback sample library, and based on the current open-source model, re-execute the steps "using the lightweight processed bird knowledge basic model, perform model fine-tuning training based on the fine-grained feature enhancement mechanism to obtain a bird knowledge multimodal model" and subsequent steps to complete the version update of the open-source model and obtain the optimized bird knowledge professional model.
[0150] Please see Figure 10 , Figure 10 This is a schematic diagram of the server structure provided in the embodiments of this application, such as... Figure 10 As shown, the server 10 includes a processor 1001, a memory 1003, a communication interface 1002, and a computer program 10031. The computer program 10031 is stored in the memory 1003 and configured to be executed by the processor 1001. The program includes a method and apparatus for training a vertical knowledge model based on an AI training and promotion integrated machine as described in the above embodiments.
[0151] This application provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implement the steps of any possible embodiment of the method.
[0152] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0153] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0154] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0155] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0156] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0157] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0158] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0159] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for training vertical knowledge models based on an AI training and promotion integrated machine, characterized in that, The method, applied to the AI training and promotion integrated machine, includes: Acquire multi-source unstructured community data of the target research subjects, including birds; The multi-source unstructured data of the community is processed to construct a multimodal training dataset for the target research object; The bird knowledge base model is lightweighted based on the multimodal training dataset. Using the lightweight bird knowledge base model, fine-tuning training is performed based on a fine-grained feature enhancement mechanism to obtain a multimodal bird knowledge model. The bird knowledge multimodal model is evaluated according to professional dimensions related to bird identification, and the model is iteratively optimized based on the evaluation results to obtain a qualified model; Based on the open-source model derived from the benchmark model, user inference requests and feedback data are received, closed-loop incremental optimization is completed, and an optimized professional model of bird knowledge is obtained.
2. The method according to claim 1, characterized in that, The process of processing the multi-source unstructured data of the community to construct the multimodal training dataset of the target research object includes: The community's multi-source unstructured data is imported, parsed, and cleaned to obtain an effective image dataset and an effective text corpus. The community's multi-source unstructured data includes bird image data originally taken by bird enthusiasts, unstructured chat log files from the community, and bird image and text materials shared within the community. The effective image dataset and the effective text corpus are subjected to multimodal image-text semantic alignment, fine-grained annotation, dataset hierarchical partitioning, and wide-area dataset expansion to obtain a multimodal training dataset.
3. The method according to claim 2, characterized in that, The process of performing multimodal image-text semantic alignment, fine-grained annotation, hierarchical partitioning of the dataset, and wide-area expansion of the dataset on the effective image dataset and the effective text corpus includes: Map the valid images in the valid image dataset to the valid text in the valid text corpus to a unified semantic space and perform semantic pairing and alignment. Automatic pre-annotation is performed on the semantically aligned image and text samples, and then the pre-annotation results are corrected through expert verification to obtain the gold standard annotated dataset; The gold standard labeled dataset is divided into a training set, a validation set, and a test set according to the principle of stratified sampling. The species distribution and the proportion of easily mixed species in each subset are consistent with the gold standard labeled dataset. The training set is fused with an open-source bird dataset, and data augmentation is performed on a small sample of species to obtain the multimodal training dataset.
4. The method according to claim 3, characterized in that, After processing the multi-source unstructured data of the community to construct the multimodal training dataset of the target research object, the method further includes: The multimodal training dataset of the AI training and push integrated machine is transmitted to the Linux computing power layer of the AI training and push integrated machine based on the shared memory mechanism.
5. The method according to any one of claims 1-4, characterized in that, The step of lightweighting the bird knowledge base model based on the multimodal training dataset includes: The bird knowledge base model is automatically matched and loaded based on the species information and sample type information of the multimodal training dataset. Based on the hardware configuration information of the AI training and promotion integrated machine, a hybrid precision quantization strategy is selected. INT4 quantization is performed on the general semantic layer of the bird knowledge base model, and INT8 quantization is performed on the bird fine-grained feature layer of the bird knowledge base model. Based on the volume of the quantized model, allocate contiguous memory and GPU memory space, and perform unified model loading and GPU memory scheduling optimization.
6. The method according to any one of claims 1-4, characterized in that, The fine-tuning training of the model based on the fine-grained feature enhancement mechanism includes: Based on the fine-grained bird recognition scenarios, corresponding training strategies are configured. The training strategies include training parameters and data augmentation strategies. The fine-grained bird recognition scenarios include scenarios for distinguishing closely related species, scenarios for recognizing individuals of different ages or sexes, scenarios for recognizing individuals under different shooting conditions, and scenarios for multimodal knowledge matching. Construct a total loss function that includes classification loss and contrastive loss, wherein the classification loss is used for basic identification and classification of bird species, and the contrastive loss is used to enhance the feature discrimination of samples of different categories; Based on the training strategy and the total loss function, the model parameters are iteratively updated using contrastive learning and hard example optimization until the preset convergence condition is met, thus completing the model fine-tuning training.
7. The method according to claim 4, characterized in that, The process of evaluating the bird knowledge multimodal model according to professional dimensions related to bird identification, and iteratively optimizing the bird knowledge multimodal model based on the evaluation results, includes: Based on the test set, inference tests are performed on the bird knowledge multimodal model to generate multi-dimensional vertical category ability evaluation indicators; The multi-dimensional vertical category ability assessment indicators are compared with preset thresholds to determine whether the bird knowledge multimodal model meets the standards. If the standard is not met, targeted optimization processing is performed according to the defect type of the non-compliant indicator, and then the steps "perform inference test on the bird knowledge multimodal model based on the test set" and subsequent steps are re-executed. If the target is met, version archiving is performed on the model that meets the target, and the model evaluation results and the target status are synchronized across systems to the Windows interaction layer.
8. The method according to claim 1, characterized in that, The process involves receiving user inference requests and feedback data from the open-source model derived from the benchmark model, completing closed-loop incremental optimization, and obtaining an optimized bird knowledge professional model, including: Receive bird identification and reasoning requests initiated by users, process the bird images in the bird identification and reasoning requests using the open-source model, and output species identification results and bird vertical category knowledge information; Collect user feedback and correction data on the recognition results, perform expert verification on the collected feedback and correction data, and filter valid feedback samples to store them in the feedback sample library; The number of valid feedback samples in the feedback sample library is counted in real time, and incremental training is triggered based on the statistical results. When incremental training is triggered, the valid feedback samples in the feedback sample library are processed, and based on the current open-source model, the steps "using the lightweight bird knowledge base model and performing model fine-tuning training based on the fine-grained feature enhancement mechanism to obtain the bird knowledge multimodal model" and subsequent steps are re-executed to complete the version update of the open-source model and obtain the optimized bird knowledge professional model.
9. A vertical knowledge model training device based on an AI training and promotion integrated machine, characterized in that, include: An acquisition unit is used to acquire multi-source unstructured social data of a target research object, including birds; The processing unit is used to process the multi-source unstructured data of the community and construct a multimodal training dataset of the target research object; The bird knowledge base model is lightweighted based on the multimodal training dataset. Using the lightweight bird knowledge foundation model, fine-tuning training is performed based on a fine-grained feature enhancement mechanism to obtain a bird knowledge multimodal model. The bird knowledge multimodal model is evaluated according to professional dimensions related to bird identification, and iterative optimization is performed on the bird knowledge multimodal model based on the evaluation results to obtain a qualified model. Based on the open-source model derived from the qualified model, user inference requests and feedback data are received to complete closed-loop incremental optimization and obtain an optimized bird knowledge professional model.
10. A server, characterized in that, It includes a processor, a memory, a communication interface, and a computer program, the computer program being stored in the memory and configured to be executed by the processor, the program including instructions for performing the steps of the method as described in any one of claims 1-8.