A text semantic label generation method, device, equipment, medium and product

CN122838565APending Publication Date: 2026-09-29JINBAOXIN SOCIAL SECURITY CARD TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611171724.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-04
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

由于非结构化文本缺乏统一的格式和规范,蕴含的信息分散且复杂,从中高效、准确地提取并生成符合要求的职位标签面临巨大挑战

Benefits of technology

本申请提供了一种文本语义标签的生成方法、装置、设备、介质及产品,通过对获取的职业培训文本数据进行预处理,能精准梳理出项目名称和项目描述等关键信息,为后续处理提供清晰有序的数据基础。基于预处理后的数据、预设任务指令以及职位细类标签集合构建设计提示信息,可以充分挖掘文本与标签间的内在联系,引导模型精准地理解文本语义。将设计提示信息输入预先构建的标签识别模型,以使模型能够依据提示信息对文本进行深度分析,从大量的职位细类标签中准确识别出目标职位细类标签,提升了职业标签提取的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838565A_ABST
    Figure CN122838565A_ABST
Patent Text Reader

Abstract

The application discloses a text semantic label generation method, device, equipment, medium and product, relates to the technical field of data processing, and comprises the following steps: preprocessing acquired vocational training text data to obtain vocational training project data; wherein the vocational training project data at least includes project name and project description information; constructing design prompt information based on the vocational training project data, and a preset task instruction and a position subcategory label set; inputting the design prompt information into a pre-constructed label recognition model to obtain a target position subcategory label output by the label recognition model, and the application improves the accuracy of vocational label extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, device, medium, and product for generating text semantic tags. Background Technology

[0002] With the continuous and in-depth advancement of social security informatization, electronic social security cards and online employment platforms are playing an increasingly important role as key infrastructure for serving users. Among these, the "vocational training voucher" service of the electronic social security card serves as a core function, providing users with a convenient and efficient way to improve their vocational skills. Users can claim and use these vouchers to participate in various online and offline vocational skills training courses, providing strong support for enhancing personal career competitiveness and promoting the healthy development of the job market.

[0003] To further optimize the connection between training and employment, detailed job category tagging of training programs has become a crucial step. Adding these tags enables subsequent intelligent job recommendations, accurately matching users with suitable employment opportunities and improving employment success rates. The accuracy of the tags determines the relevance and usability of the intelligent job recommendation results, which is vital to improving the effectiveness of the employment service system.

[0004] However, in practice, it has been found that the system generates massive amounts of training program data daily. This data mainly exists in the form of unstructured text. Due to the lack of a unified format and standards for unstructured text, the information it contains is scattered and complex, posing a significant challenge to efficiently and accurately extracting and generating compliant job tags. Summary of the Invention

[0005] The purpose of this application is to provide a method, apparatus, device, medium, and product for generating text semantic tags, which can improve the accuracy of occupation tag extraction.

[0006] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for generating text semantic tags, including: The acquired vocational training text data is preprocessed to obtain vocational training project data; wherein, the vocational training project data includes at least project name and project description information; Based on the vocational training project data and the preset task instructions and job category tag set, design prompt information is constructed; The design prompts are input into a pre-built tag recognition model to obtain the target job category tags output by the tag recognition model.

[0007] Optionally, the preprocessing of the acquired vocational training text data to obtain vocational training project data specifically includes: Determine the acquisition time and location corresponding to the obtained vocational training text data; Determine the computing node that matches the acquisition time and the acquisition location; The target vocational training text data is obtained by preprocessing the vocational training text data through the computing nodes. Semantic recognition is performed on the target vocational training text data to obtain the project name and project description information; The project name and project description information are identified as vocational training project data.

[0008] Optionally, the step of preprocessing the vocational training text data through the computing node to obtain the target vocational training text data specifically includes: Identify special characters in the vocational training text data; wherein the character type of the special characters is not text. The special characters in the vocational training text data are deleted to obtain the current vocational training text data; Each character in the current vocational training text data is formatted to obtain the target vocational training text data; wherein, each character in the target vocational training text data has the same character format.

[0009] Optionally, the step of constructing design prompt information based on the vocational training project data and a preset set of task instructions and job category tags specifically includes: Obtain a preset set of task instructions and job category tags; wherein, the set of job category tags contains multiple pre-added job category tags; each job category tag corresponds to a job sub-category tag, a job major category tag, and a job name; Obtain output format information; Design prompts are constructed based on the vocational training project data, the task instructions, the job category tag set, and the output format information.

[0010] Optionally, the label recognition model is constructed as follows: Determine the target basic model; The target base model is pre-trained to obtain an initial label recognition model; wherein, the initial label recognition model is used to output training job category labels corresponding to the training design prompts input to the initial label recognition model; The initial tag recognition model is deployed to a server cluster to obtain a tag recognition model; wherein, the tag recognition model provides an API interface, the tag recognition model is connected to a database, and the database stores the set of job category tags.

[0011] Optionally, after obtaining the target job category label output by the label recognition model, the method further includes: Obtain the geographic location and career preference information of the user who inputs the vocational training text data; Get online job information; From the online job information, identify target job information that matches the geographical location, the career preference information, and the target job category tags; Output the target job information.

[0012] Secondly, this application provides a text semantic tag generation apparatus, comprising: The preprocessing unit is used to preprocess the acquired vocational training text data to obtain vocational training project data; wherein, the vocational training project data includes at least project name and project description information; The construction unit is used to construct design prompt information based on the vocational training project data and the preset set of task instructions and job category tags; The input unit is used to input the design prompt information into a pre-built label recognition model to obtain the target job category label output by the label recognition model.

[0013] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the text semantic tag generation method described in any one of the above.

[0014] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the text semantic tag generation method described above.

[0015] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the text semantic tag generation method described above.

[0016] In a sixth aspect, this application provides a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run a program or instructions, and the processor executing the program or instructions implementing the steps of the text semantic tag generation method described above.

[0017] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method, apparatus, device, medium, and product for generating text semantic tags. By preprocessing the acquired vocational training text data, key information such as project names and descriptions can be accurately extracted, providing a clear and orderly data foundation for subsequent processing. Design prompts are constructed based on the preprocessed data, preset task instructions, and a set of job category tags. This fully explores the inherent relationship between text and tags, guiding the model to accurately understand the text semantics. The design prompts are input into a pre-built tag recognition model, enabling the model to perform in-depth analysis of the text based on the prompts, accurately identifying the target job category tag from a large number of job category tags, thus improving the accuracy of job tag extraction. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a method for generating text semantic tags according to an embodiment of this application; Figure 2 A schematic diagram of the functional modules of a text semantic tag generation device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] In one exemplary embodiment, such as Figure 1 As shown, a method for generating text semantic tags is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, it includes steps 101 to 103. Wherein: Step 101: Preprocess the obtained vocational training text data to obtain vocational training project data.

[0023] In this embodiment of the application, the vocational training project data includes at least the project name and project description information.

[0024] For example, the project name could be Level 4 Public Nutritionist, and the project description could be "The training course includes basic nutrition, food nutrition evaluation, and dietary guidance for chronic diseases. For example, practical skills such as diabetes diet design and weight management methods."

[0025] As an optional implementation, step 101 preprocesses the acquired vocational training text data to obtain vocational training project data, which may include: Determine the acquisition time and location corresponding to the obtained vocational training text data; Determine the computing node that matches the acquisition time and the acquisition location; The target vocational training text data is obtained by preprocessing the vocational training text data through the computing nodes. Semantic recognition is performed on the target vocational training text data to obtain the project name and project description information; The project name and project description information are identified as vocational training project data.

[0026] This implementation method, by matching computing nodes based on acquisition time and location, can rationally allocate computing resources, improve processing efficiency, and reduce data transmission latency. Preprocessing the vocational training text data through the matched computing nodes allows for targeted optimization of the processing flow, resulting in more accurate target data. Semantic recognition of the target data accurately extracts project names and descriptions, identifying it as vocational training project data. This provides a high-quality, structured data foundation for subsequent operations, contributing to improved accuracy and stability of the entire text semantic tag generation process.

[0027] In this embodiment, a matching computing node is determined based on the acquisition time and location of the vocational training text data. Subsequent operations can be performed through this computing node, reducing the amount of data in a single task and improving processing efficiency.

[0028] Optionally, the method of preprocessing the vocational training text data through the computing node to obtain the target vocational training text data may include: Identify special characters in the vocational training text data; wherein the character type of the special characters is not text. The special characters in the vocational training text data are deleted to obtain the current vocational training text data; Each character in the current vocational training text data is formatted to obtain the target vocational training text data; wherein, each character in the target vocational training text data has the same character format.

[0029] This implementation method identifies and removes non-textual special characters from vocational training text data, eliminating interfering elements and preventing them from negatively impacting subsequent processing, thus improving data purity. A standardized format conversion is then performed on the vocational training text data after removing special characters, ensuring consistent character formatting and enhancing data standardization and compatibility. This not only facilitates unified data processing and analysis but also effectively reduces errors caused by format differences, laying a solid foundation for generating accurate text semantic tags and improving the efficiency and accuracy of the entire processing flow.

[0030] In this embodiment of the application, special characters can be irrelevant symbols (such as "#", "*"), punctuation marks (such as redundant ",", ".", etc.).

[0031] In this embodiment of the application, the format conversion of each character in the current vocational training text data can be performed by converting each character in the current vocational training text data into a unified encoding format (such as UTF-8) and standardizing the case, spaces, etc.

[0032] In this embodiment, the final target vocational training text data can be stored in the database in the format of a pre-set vocational training project table. This database can be a Hive database from a big data platform.

[0033] In this embodiment of the application, the format of the vocational training project table in the database can be: Project ID; Project Name; Project description.

[0034] For example: Project ID: TRNxxxx0001; Project Name: Level 4 Public Nutritionist; Project Description: "The training course includes topics such as basic nutrition, food nutrition evaluation, and dietary guidance for chronic diseases. Practical skills include designing diabetic diets and weight management methods."

[0035] Step 102: Construct design prompt information based on the vocational training project data and the preset task instructions and job category tag set.

[0036] In this embodiment, the design prompt can be used to guide the generation of target job category tags that meet the requirements. The design prompt needs to conform to the request format required by the interface of the design prompt; specifically, it can be: ①Task instructions: Clearly define the task objectives, which can be fixed content, such as "Analyze the job category tags corresponding to the project based on the project name and project description".

[0037] For example, this section contains the core task instruction of the prompt, clearly telling the label recognition model the target it needs to perform. This section remains fixed and is specifically stated as: "Based on the project name and description, analyze the corresponding job category label for this project, and output only the most suitable label." This instruction ensures that the model focuses on the label classification task and restricts the output to a unique label, avoiding the generation of redundant content.

[0038] ② Job category tag set: This serves as the candidate range for the tag recognition model to generate tags. It can be fixed content. For example, the job category tag set can contain 1635 job category tags.

[0039] For example: edible fungi production; tropical crop cultivation; Chinese medicinal herb growers; forestry and grassland seedling workers; afforestation and regeneration workers; forest rangers; forest tending workers; timber harvesting workers; timber water transport workers; livestock breeders; poultry breeders; livestock feeders; poultry feeders; economic insect breeders; ... (a total of 1635 items); ③ Vocational training project data: The project name and project description are used as inputs and provided to the label recognition model.

[0040] For example: Project Name: Livestock Farmer Training; Project Description: Training in the basic theories and practical skills of livestock breeding.

[0041] This section will be replaced with the actual data.

[0042] ④ Output format information: Explicitly require the tag recognition model to output only the most suitable job category tag and return it in a preset output format (e.g., JSON format), which can be fixed content.

[0043] For example, the output format information can be expressed as: "Only output the most suitable tag. The output format is JSON. Example: { "result": "tag name"}.

[0044] As an optional implementation, step 102, which constructs design prompt information based on the vocational training project data and a preset set of task instructions and job category tags, may include: Obtain a preset set of task instructions and job category tags; wherein, the set of job category tags contains multiple pre-added job category tags; each job category tag corresponds to a job sub-category tag, a job major category tag, and a job name; Obtain output format information; Design prompts are constructed based on the vocational training project data, the task instructions, the job category tag set, and the output format information.

[0045] This implementation method, by acquiring a set of job tagging details containing multiple job subcategories and their associated job classes, major categories, and names, provides comprehensive and detailed job tagging information, laying the foundation for accurate matching. Obtaining output format information clarifies the result presentation requirements. Constructing design prompts based on vocational training project data, task instructions, tag sets, and output format information allows these prompts to cover multiple key aspects, guiding subsequent processing to more accurately understand the requirements. It accurately filters and generates suitable semantic tags from a wealth of information, effectively improving the accuracy and relevance of tag generation and enhancing the practicality and effectiveness of the entire process.

[0046] In this embodiment of the application, each job category tag in the job category tag set can be pre-stored in the database in the format of a job data table, wherein the specific format of the job data table can be: Job ID; Job title; Job categories; Job Category; Job categories.

[0047] For example: Job ID: 143xxxxxx9; Job Title: Public Nutritionist; Job Category: Health Services; Job Category: Health Consultant Service Personnel; Job category: Nutritionist.

[0048] Step 103: Input the design prompt information into the pre-built tag recognition model to obtain the target job category tag output by the tag recognition model.

[0049] As an optional implementation method, the label recognition model can be constructed in the following ways: Determine the target basic model; The target base model is pre-trained to obtain an initial label recognition model; wherein, the initial label recognition model is used to output training job category labels corresponding to the training design prompts input to the initial label recognition model; The initial tag recognition model is deployed to a server cluster to obtain a tag recognition model; wherein, the tag recognition model provides an API interface, the tag recognition model is connected to a database, and the database stores the set of job category tags.

[0050] This implementation method, by identifying and pre-training the target base model, enables the model to initially possess recognition capabilities, generating an initial label recognition model capable of outputting training job category labels, thus laying the foundation for accurate recognition. Deploying it to a server cluster fully utilizes the cluster's powerful computing capabilities, improving model processing efficiency and stability. Providing an API interface for easy access and connection to a database storing the set of job category labels allows the model to quickly acquire rich label information, enabling more accurate matching and output of target job category labels when processing text, effectively improving the accuracy and efficiency of text semantic label generation, and enhancing the system's practicality and scalability.

[0051] In this embodiment, the target base model can be deepseek-r1:32b. Pre-training the target base model yields an initial label recognition model with strong semantic understanding and support for Chinese, making it suitable for processing complex text data from vocational training projects.

[0052] Furthermore, pre-trained deepseek-r1:32b models can be loaded onto the server cluster. The model is then encapsulated as a RESTful service using the API provided by Ollam, supporting HTTP requests. The model service is deployed on a separate server cluster and connected to the Hive cluster via a network. The service address (e.g., http: / / xxxx:11434 / api / generate) is provided to the Hive UDF (User-Defined Function). The UDF internally encapsulates the logic for calling the model interface, mapping text data such as project names and descriptions to job category tags. The UDF is the core component of this invention for achieving deep integration between the big data platform and the model.

[0053] Specifically, the UDF code is packaged into a JAR file, uploaded to HDFS (Hadoop Distributed File System), and registered in the Hive environment. After registration, the UDF can be directly called in the Hive database.

[0054] Optionally, when processing data from multiple vocational training projects, the deployed UDF can be invoked through the Hive database to batch process the data from multiple vocational training projects in the Hive database and generate tags.

[0055] Specifically: 1. Task Distribution: The Hive execution engine (MapReduce, Tez, or Spark) distributes data from multiple vocational training projects to various computing nodes.

[0056] 2. Parallel processing: Each computing node executes UDF code in parallel, sends requests to the label recognition model, and processes the returned results.

[0057] During Hive task execution, multiple vocational training project data are distributed to multiple computing nodes, and project records are processed in parallel row by row. Each computing node reads the vocational training project data one by one according to the storage order. For each vocational training project data, the computing node extracts the project name and project description, inserts them as variable parameters into a preset prompt template, and uses this to construct design prompt information for the tag recognition model. After construction, the computing node sends the design prompt information to the deployed tag recognition model to obtain the target job category tags generated by the tag recognition model. The generated target job category tags are then bound to the original records for subsequent result aggregation and writing operations.

[0058] 3. Result aggregation: The label results generated by each computing node are aggregated into the Hive database.

[0059] The generated target job category tags are then bound to the original records to obtain the target job information table, which can be structured as follows: Project ID (project_id); Project name (project_name); Project description (project_description); Job category tags.

[0060] Furthermore, the target job information table can be stored in a Hive database for later use.

[0061] In this embodiment, design hints can be sent to the deployed tag recognition model via a POST request using the HTTP protocol. The UDF internally encapsulates an HTTP client (such as Java's HttpClient) responsible for communicating with the tag recognition model.

[0062] The system receives the target job category tags (i.e., the values ​​of the result field) returned by the tag recognition model. The parsed tags are then returned to the Hive execution engine as the output of the UDF.

[0063] To ensure the stability of compute nodes, UDF implements a robust exception handling mechanism, including: network anomalies, request errors, and logging. Specifically: Network error: Handles network timeouts, service unavailability, and other situations, triggering a retry mechanism (e.g., retries up to 3 times). Request Error: Handles cases where the request format is incorrect or the tag recognition model returns an error, logs the error, and returns a default value (such as "unknown tag"). Log recording: Detailed log recording of input, output and exceptions for each request, facilitating subsequent debugging and optimization.

[0064] As an optional implementation, after step 103, the following steps may also be performed: Obtain the geographic location and career preference information of the user who inputs the vocational training text data; Get online job information; From the online job information, identify target job information that matches the geographical location, the career preference information, and the target job category tags; Output the target job information.

[0065] This implementation method, by acquiring users' geographical location and career preference information, allows for a precise understanding of users' actual needs and environment. Combined with pre-generated target job category tags, matching target job information is filtered from online job listings, enabling precise and personalized job recommendations. This multi-dimensional matching approach considers both users' career inclinations and geographical location factors, pushing more realistic and practical jobs to users, significantly increasing their chances of finding suitable employment, enhancing user experience, and improving the value and practicality of the entire vocational training and employment service system.

[0066] The core innovation of this application lies in the deep integration of Hive UDF and large model services. Through Hive UDF, seamless integration of the big data platform (Hive) and large model inference capabilities is achieved. Traditional solutions typically require exporting data from Hive to external systems for processing. However, this invention directly calls large model services within Hive via UDF, fully leveraging Hive's distributed computing capabilities and the powerful semantic understanding capabilities of pre-trained large models. This combination is key to achieving efficient and accurate processing of massive amounts of data.

[0067] Prompt design is a key factor in guiding large models to generate job category labels that meet the requirements. Different prompt designs may lead to different generation results. However, this invention, through a carefully designed prompt, makes full use of the semantic understanding capabilities of large models and solves problems such as semantic ambiguity and expression diversity.

[0068] The structure and content of a Prompt include the instruction section, the context section (a set of job category tags), the input data section, and the design method for output format requirements.

[0069] How the candidate tag set is organized: The definition and organization of the 1635 job category tags in the Prompt, for example, in the form of a semicolon-separated list.

[0070] Prompt's dynamic update mechanism supports adding, deleting, and modifying job category tag sets, such as dynamically loading tag sets through configuration files or databases.

[0071] This application fully leverages Hive's distributed execution engine (such as MapReduce, Tez, or Spark) to achieve distributed parallel processing of label generation tasks. Each computing node independently calls the UDF, sends requests to the large model service, and processes the returned results, thereby efficiently processing massive amounts of data.

[0072] Distributed task distribution mechanism: The way the Hive execution engine distributes data to various computing nodes, including partitioning and parallelism adjustment.

[0073] The execution logic of UDF in a distributed environment: the parallel invocation method of UDF on each computing node, including request distribution, result aggregation, etc.

[0074] Performance optimization strategies include partitioning, parallelism adjustment, and other optimization techniques.

[0075] This application designs a structured storage method for tagging results and seamlessly integrates it into business applications. The generated tags are stored in a Hive table, forming a standardized data structure that facilitates subsequent applications such as job recommendations and data analysis.

[0076] The storage structure of the tagged results: The field design of the Hive table tagged_training_projects includes project ID, project name, project description, job category tags, etc.

[0077] Tag result retrieval mechanism: How tag data is retrieved in business applications such as job recommendation and data analysis, for example, through SQL queries.

[0078] Application scenarios for tag results include specific applications such as intelligent job recommendations.

[0079] Although this application adopts the technical solution of Hive UDF + DeepSeek + Ollama, the following alternatives can also achieve similar purposes: 1. Custom MapReduce / Tez tasks + large model service Solution Description: Directly write MapReduce or Tez tasks to call the large model service during the Map or Reduce phase. Specifically, during the Map phase, read data from the Hive table, send requests to the large model service to generate tags, and during the Reduce phase, summarize the results and write them to the Hive table.

[0080] Advantages: High flexibility, allowing customization of MapReduce or Tez task logic to meet specific needs. Supports complex distributed processing logic, such as multi-stage label generation or post-processing.

[0081] Compared with the present invention: This application achieves seamless integration with Hive SQL through UDF. Users only need to write simple SQL statements to complete the tag generation task, resulting in lower development and maintenance costs, while the performance is comparable to the MapReduce / Tez solution.

[0082] 2. Standalone programs written in Python / Java / other languages ​​+ large model services Solution Description: Write independent programs using languages ​​such as Python and Java to read data from Hive, call the large model service to generate tags, and write the results back to Hive. Specifically, the program reads data from Hive via JDBC or Hive CLI, calls the API interface of the large model service, and finally writes the results back to the Hive table using SQL insert statements.

[0083] Advantages: Highest flexibility, capable of implementing complex business logic, such as multi-model integration and post-processing rules. Diverse development language options, suitable for different development teams' technology stacks.

[0084] Compared with the present invention: This application generates tags directly within Hive using UDFs, without the need for data migration or independent program development, resulting in better performance and integration, and lower development and maintenance costs.

[0085] 3. Use other large model services Solution Description: Replace DeepSeek with other large models, such as ChatGLM, Baichuan, LLaMA, or OpenAI's GPT series models. Specifically, replace the API interface of the large model service with the interface of the other model, while keeping the interaction logic between the UDF and the large model service unchanged.

[0086] Advantages: A variety of models are available, allowing users to choose the most suitable model based on performance, cost, or support for Chinese. Replacement costs are low; the UDF design of this invention is modular, requiring only modifications to the interface address and request format to adapt to other models.

[0087] In comparison with this invention: This application chose the DeepSeek model because of its strong support for Chinese characters, low deployment cost, and good compatibility with the Ollam framework. Although other models are theoretically feasible, the UDF design of this invention fully considers the model's versatility and supports rapid adaptation to other models, thus offering greater advantages in practical applications.

[0088] While the above alternatives are technically feasible, the advantages of the proposed solution are as follows: 1. Make full use of Hive infrastructure: Through UDF and Hive's distributed execution engine, efficient parallel processing of tag generation is achieved without the need for additional middleware or data migration.

[0089] 2. Deep integration of large models: By leveraging the powerful semantic understanding capabilities of pre-trained large models, problems such as semantic ambiguity and expression diversity are solved, while reducing the cost of manual annotation and model training.

[0090] 3. Low development and maintenance costs: The modular design of UDF and the dynamic update mechanism of Prompt make the system easy to develop, deploy and maintain.

[0091] 4. High scalability: Supports the replacement of different large models or computing engines, and adapts to the dynamic changes of job category tags.

[0092] In summary, the technical solution of this application is superior to the above-mentioned alternatives in terms of performance, cost, accuracy, integration, and scalability, and is the best choice to achieve the purpose of the invention.

[0093] It is evident that this application can improve processing efficiency: by calling large models through Hive User-Defined Functions (UDFs), distributed parallel processing of tag generation is achieved, making full use of Hive's distributed computing capabilities to meet the needs of efficient processing of massive amounts of data.

[0094] This application can also reduce development and maintenance costs: by leveraging the general semantic understanding capabilities of pre-trained large models, it reduces reliance on manually labeled data and model training, thereby lowering development and maintenance costs.

[0095] This application can also improve label accuracy: by leveraging the superior semantic understanding capabilities of large models, it solves problems such as semantic ambiguity, expression diversity, and the generation of labels for emerging job titles, ensuring the accuracy and coverage of label generation.

[0096] This application also enables native big data integration: without intermediate data migration or additional middleware development, data processing and tag generation can be completed directly in Hive, reducing system complexity and integration costs.

[0097] This application also enhances scalability and adaptability: it designs a flexible system architecture that supports the replacement of different large models or computing engines and automatically adapts to dynamic changes in job categories (such as the addition of new job categories) without requiring significant adjustments to the system architecture.

[0098] Implementing steps 101 to 103 above enables accurate identification of target job category tags from a large number of job category tags, improving the accuracy of job tag extraction. Furthermore, this application provides a high-quality, structured data foundation for subsequent operations, contributing to improved accuracy and stability of the entire text semantic tag generation process. In addition, this application facilitates unified data processing and analysis, effectively reducing errors caused by format differences, laying a solid foundation for generating accurate text semantic tags, and improving the efficiency and accuracy of the entire processing flow. Furthermore, this application effectively improves the accuracy and relevance of tag generation, enhancing the practicality and effectiveness of the entire process. Moreover, this application effectively improves the accuracy and efficiency of text semantic tag generation, enhancing the system's practicality and scalability. Furthermore, this application can push more realistic and practical job postings to users, greatly increasing their chances of finding suitable employment, enhancing user experience, and improving the value and practicality of the entire vocational training and employment service system.

[0099] Based on the same inventive concept, this application also provides a text semantic tag generation apparatus for implementing the text semantic tag generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations of one or more text semantic tag generation apparatus embodiments provided below can be found in the limitations of the text semantic tag generation method described above, and will not be repeated here.

[0100] In one exemplary embodiment, such as Figure 2 As shown, a text semantic tag generation device is provided, comprising: The preprocessing unit 201 is used to preprocess the acquired vocational training text data to obtain vocational training project data; wherein, the vocational training project data includes at least project name and project description information; Construction unit 202 is used to construct design prompt information based on the vocational training project data and a preset set of task instructions and job category tags; The input unit 203 is used to input the design prompt information into the pre-built label recognition model to obtain the target job category label output by the label recognition model.

[0101] As an optional implementation, the preprocessing unit 201 preprocesses the acquired vocational training text data to obtain vocational training project data in the following specific ways: Determine the acquisition time and location corresponding to the obtained vocational training text data; Determine the computing node that matches the acquisition time and the acquisition location; The target vocational training text data is obtained by preprocessing the vocational training text data through the computing nodes. Semantic recognition is performed on the target vocational training text data to obtain the project name and project description information; The project name and project description information are identified as vocational training project data.

[0102] This implementation method, by matching computing nodes based on acquisition time and location, can rationally allocate computing resources, improve processing efficiency, and reduce data transmission latency. Preprocessing the vocational training text data through the matched computing nodes allows for targeted optimization of the processing flow, resulting in more accurate target data. Semantic recognition of the target data accurately extracts project names and descriptions, identifying it as vocational training project data. This provides a high-quality, structured data foundation for subsequent operations, contributing to improved accuracy and stability of the entire text semantic tag generation process.

[0103] As an optional implementation, the preprocessing unit 201 preprocesses the vocational training text data through the computing node to obtain the target vocational training text data in the following specific ways: Identify special characters in the vocational training text data; wherein the character type of the special characters is not text. The special characters in the vocational training text data are deleted to obtain the current vocational training text data; Each character in the current vocational training text data is formatted to obtain the target vocational training text data; wherein, each character in the target vocational training text data has the same character format.

[0104] This implementation method identifies and removes non-textual special characters from vocational training text data, eliminating interfering elements and preventing them from negatively impacting subsequent processing, thus improving data purity. A standardized format conversion is then performed on the vocational training text data after removing special characters, ensuring consistent character formatting and enhancing data standardization and compatibility. This not only facilitates unified data processing and analysis but also effectively reduces errors caused by format differences, laying a solid foundation for generating accurate text semantic tags and improving the efficiency and accuracy of the entire processing flow.

[0105] As an optional implementation, the construction unit 202 constructs design prompt information based on the vocational training project data and a preset set of task instructions and job category tags in the following specific ways: Obtain a preset set of task instructions and job category tags; wherein, the set of job category tags contains multiple pre-added job category tags; each job category tag corresponds to a job sub-category tag, a job major category tag, and a job name; Obtain output format information; Design prompts are constructed based on the vocational training project data, the task instructions, the job category tag set, and the output format information.

[0106] This implementation method, by acquiring a set of job tagging details containing multiple job subcategories and their associated job classes, major categories, and names, provides comprehensive and detailed job tagging information, laying the foundation for accurate matching. Obtaining output format information clarifies the result presentation requirements. Constructing design prompts based on vocational training project data, task instructions, tag sets, and output format information allows these prompts to cover multiple key aspects, guiding subsequent processing to more accurately understand the requirements. It accurately filters and generates suitable semantic tags from a wealth of information, effectively improving the accuracy and relevance of tag generation and enhancing the practicality and effectiveness of the entire process.

[0107] As an optional implementation method, the label recognition model can be constructed in the following ways: Determine the target basic model; The target base model is pre-trained to obtain an initial label recognition model; wherein, the initial label recognition model is used to output training job category labels corresponding to the training design prompts input to the initial label recognition model; The initial tag recognition model is deployed to a server cluster to obtain a tag recognition model; wherein, the tag recognition model provides an API interface, the tag recognition model is connected to a database, and the database stores the set of job category tags.

[0108] This implementation method, by identifying and pre-training the target base model, enables the model to initially possess recognition capabilities, generating an initial label recognition model capable of outputting training job category labels, thus laying the foundation for accurate recognition. Deploying it to a server cluster fully utilizes the cluster's powerful computing capabilities, improving model processing efficiency and stability. Providing an API interface for easy access and connection to a database storing the set of job category labels allows the model to quickly acquire rich label information, enabling more accurate matching and output of target job category labels when processing text, effectively improving the accuracy and efficiency of text semantic label generation, and enhancing the system's practicality and scalability.

[0109] As an optional implementation, the input unit 203 is also used for: After obtaining the target job category labels output by the label recognition model, the geographic location and job preference information of the user who input the vocational training text data are obtained; Get online job information; From the online job information, identify target job information that matches the geographical location, the career preference information, and the target job category tags; Output the target job information.

[0110] This implementation method, by acquiring users' geographical location and career preference information, allows for a precise understanding of users' actual needs and environment. Combined with pre-generated target job category tags, matching target job information is filtered from online job listings, enabling precise and personalized job recommendations. This multi-dimensional matching approach considers both users' career inclinations and geographical location factors, pushing more realistic and practical jobs to users, significantly increasing their chances of finding suitable employment, enhancing user experience, and improving the value and practicality of the entire vocational training and employment service system.

[0111] Implementing the above-described methods enables accurate identification of target job category tags from a large number of job category tags, improving the accuracy of job tag extraction. Furthermore, this application provides a high-quality, structured data foundation for subsequent operations, contributing to improved accuracy and stability of the entire text semantic tag generation process. In addition, this application facilitates unified data processing and analysis, effectively reducing errors caused by format differences, laying a solid foundation for generating accurate text semantic tags, and improving the efficiency and accuracy of the entire processing flow. Furthermore, this application effectively improves the accuracy and relevance of tag generation, enhancing the practicality and effectiveness of the entire process. Moreover, this application effectively improves the accuracy and efficiency of text semantic tag generation, enhancing the system's practicality and scalability. Furthermore, this application can push more realistic and practical job postings to users, significantly increasing their chances of finding suitable employment, enhancing user experience, and improving the value and practicality of the entire vocational training and employment service system.

[0112] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 3As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data for generating text semantic tags. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for generating text semantic tags.

[0113] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0114] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0115] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0116] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0117] In one exemplary embodiment, a chip is provided, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps in the above method embodiments and achieve the same technical effect, and will not be described again here to avoid repetition.

[0118] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0120] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0121] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0123] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for generating text semantic tags, characterized in that, The method for generating the text semantic tags includes: The acquired vocational training text data is preprocessed to obtain vocational training project data; wherein, the vocational training project data includes at least project name and project description information; Based on the vocational training project data and the preset task instructions and job category tag set, design prompt information is constructed; The design prompts are input into a pre-built tag recognition model to obtain the target job category tags output by the tag recognition model.

2. The method for generating text semantic tags according to claim 1, characterized in that, The preprocessing of the acquired vocational training text data to obtain vocational training project data specifically includes: Determine the acquisition time and location corresponding to the obtained vocational training text data; Determine the computing node that matches the acquisition time and the acquisition location; The target vocational training text data is obtained by preprocessing the vocational training text data through the computing nodes. Semantic recognition is performed on the target vocational training text data to obtain the project name and project description information; The project name and project description information are identified as vocational training project data.

3. The method for generating text semantic tags according to claim 2, characterized in that, The step of preprocessing the vocational training text data through the computing node to obtain the target vocational training text data specifically includes: Identify special characters in the vocational training text data; wherein the character type of the special characters is not text. The special characters in the vocational training text data are deleted to obtain the current vocational training text data; Each character in the current vocational training text data is formatted to obtain the target vocational training text data; wherein, each character in the target vocational training text data has the same character format.

4. The method for generating text semantic tags according to claim 1, characterized in that, The construction of design prompt information based on the vocational training project data and the preset set of task instructions and job category tags specifically includes: Obtain a preset set of task instructions and job category tags; wherein, the set of job category tags contains multiple pre-added job category tags; each job category tag corresponds to a job sub-category tag, a job major category tag, and a job name; Obtain output format information; Design prompts are constructed based on the vocational training project data, the task instructions, the job category tag set, and the output format information.

5. The method for generating text semantic tags according to claim 1, characterized in that, The specific construction method of the label recognition model is as follows: Determine the target basic model; The target base model is pre-trained to obtain an initial label recognition model; wherein, the initial label recognition model is used to output training job category labels corresponding to the training design prompts input to the initial label recognition model; The initial tag recognition model is deployed to a server cluster to obtain a tag recognition model; wherein, the tag recognition model provides an API interface, the tag recognition model is connected to a database, and the database stores the set of job category tags.

6. The method for generating text semantic tags according to claim 1, characterized in that, After obtaining the target job category label output by the label recognition model, the method further includes: Obtain the geographic location and career preference information of the user who inputs the vocational training text data; Get online job information; From the online job information, identify target job information that matches the geographical location, the career preference information, and the target job category tags; Output the target job information.

7. A device for generating text semantic tags, characterized in that, The device for generating text semantic tags includes: The preprocessing unit is used to preprocess the acquired vocational training text data to obtain vocational training project data; wherein, the vocational training project data includes at least project name and project description information; The construction unit is used to construct design prompt information based on the vocational training project data and the preset set of task instructions and job category tags; The input unit is used to input the design prompt information into a pre-built label recognition model to obtain the target job category label output by the label recognition model.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the method for generating text semantic tags according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method for generating text semantic tags as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method for generating text semantic tags as described in any one of claims 1-6.