Recruitment information supervision platform and supervision method based on multi-channel integration
By processing key fields of recruitment information, aggregating information from the same source across channels, and assessing risks in multiple dimensions, this technology addresses the problem of insufficient information detail depiction in existing cross-channel recruitment information management, and achieves efficient supervision and risk identification of cross-channel recruitment information.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-03
AI Technical Summary
Existing multi-channel recruitment information management technologies lack sufficient support in cross-channel semantic association mining and homogeneous information merging and integration, time-series information version evolution tracking, and rule matching. They struggle to identify the same recruiting entity's disguised posting behavior on different platforms and frequent modifications to job descriptions, resulting in insufficient identification of abnormal recruitment information, increased risk assessment bias, longer operation chains, and higher costs.
The key field processing module extracts recruitment information metadata for standardized processing, generates standardized job records and performs five-tuple modeling, the copywriting evolution generation module performs cross-channel homogeneous aggregation to construct a recruitment information graph, the tampering risk aggregation module calculates the differences between adjacent versions to perform multi-dimensional risk judgment, and the regulatory process iteration module performs adaptive iterative optimization.
It has improved cross-channel consistency inspection capabilities, enhanced dynamic parameter adjustment capabilities, and improved the ability to identify complex tampering behaviors and hidden differences across channels, reducing the pressure of manual review and improving the stability of the regulatory closed loop and the accuracy of risk identification.
Smart Images

Figure CN121788086A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data processing technology, specifically to a recruitment information supervision platform and method based on multi-channel integration. Background Technology
[0002] Existing multi-channel information management follows a standardized closed-loop process: First, stable data connections are established with various information source systems through diverse methods such as interface integration and page crawling, and raw information scattered across different channels is collected in batches at preset time intervals. Then, in the background, the core fields of the collected data are parsed, redundancy is cleaned, and standardized, and heterogeneous data is uniformly mapped to preset field templates and stored in a centralized information database. On this basis, the configured keyword database, blacklist database, and rule scripts are called to automatically verify newly added or changed information, and screen out suspected violations that violate prohibited clauses or match risk characteristics. Finally, the suspected violations are pushed to the manual review stage through the management interface, where reviewers verify the original information carrier, supplement judgment opinions, and record the processing results. In some scenarios, basic statistical reports are also generated based on manually processed data to provide a reference for optimizing rule configuration and expanding the risk database.
[0003] This management process has been specifically applied in the field of recruitment information management. Because recruitment information is scattered across multiple channels such as mainstream recruitment websites, third-party human resource service systems, and enterprise-owned recruitment systems, and needs to meet specific requirements such as compliance supervision and information authenticity verification, the standardized process for multi-channel information management has been further adapted. General information is focused on recruitment information, core fields are clearly defined as content specific to the recruitment scenario, and rule verification is optimized around the compliance requirements of the recruitment industry, ultimately forming a specialized management model for multi-channel recruitment information.
[0004] For example, Chinese invention patent application CN119988580A discloses a method, apparatus, and medium for managing human resource text information based on big data. The method includes: acquiring text data of the object to be analyzed collected from multiple channels; extracting representative features from the text data; processing the representative features to transform them into features of another dimension; classifying and grouping the processed representative features; and calculating the feature score of the text data of the object to be analyzed based on the results of feature grouping and feature classification.
[0005] For example, Chinese invention patent application CN117725098A discloses a method and electronic device for managing massive user multi-channel data, belonging to the field of data processing technology. The method involves: retrieving a cached database based on the channel number to obtain a channel template; based on the channel template, including fixed routes and custom routes, including: performing a first mod operation on the account ID according to the first fixed route to obtain fixed database sharding information; performing a second mod operation on the ID according to the second fixed route to obtain account sharding information, querying the sharding information corresponding to the account sharding information to obtain the user's uid; performing a third mod operation on the uid according to the third fixed route to obtain user sharding information; performing a mod operation based on the first custom route and the account ID to obtain custom database sharding information; performing a mod operation based on the second custom route and the uid to obtain custom sharding information; and returning the queried data information as the result.
[0006] Most existing multi-channel information management solutions are based on a single-record, single-channel, static text processing approach. They project multi-source data into a single feature space for batch statistics, or mainly rely on routing templates and database sharding mechanisms to solve multi-channel access and concurrency issues. In terms of architecture design, they focus more on optimizing the data entry process for single records and calculating static scores. In terms of cross-channel semantic association mining and integration of information from the same source, tracking the evolution of information versions in the time series dimension, and full-link traceability of rule-hitting processes, current solutions need to strengthen their support for multi-dimensional association scenarios. Their ability to characterize information details is still insufficient, and their flexibility in dynamic evolution adaptation and continuous optimization needs to be improved. Overall, they show that there is room for improvement in the coverage of technical support, the refinement of detail characterization, and dynamic adaptation capabilities.
[0007] In general business scenarios that emphasize data aggregation and statistical analysis, this type of design can still meet the basic requirements of being able to store data, retrieve information, and perform calculations quickly. However, in scenarios that require multi-perspective comparisons, behavioral pattern decomposition, and long-term evolution characterization based on the text content itself, the above features will gradually become limiting factors.
[0008] Multi-channel recruitment information management is a typical scenario where the consistency of job descriptions across different channels, the behavioral patterns of the same entity under different accounts, and the modification and reposting of job descriptions over time are particularly sensitive. In this scenario, these technical shortcomings are amplified: on the one hand, the system struggles to form a stable cross-platform, cross-version aggregation perspective when the same recruiting entity posts under different names on different platforms, or when the description of the same job posting differs across channels; on the other hand, it is not sensitive enough to the time evolution patterns of frequent modifications, deletions, and reposts of job descriptions. When it is necessary to reconstruct the page content, judgment criteria, and timeline afterward, the operation chain is long and costly, resulting in a series of negative impacts, including insufficient identification of abnormal recruitment information, increased risk assessment bias, and a passively prolonged processing loop. Summary of the Invention
[0009] In view of the shortcomings of the prior art, the present invention provides a recruitment information supervision platform and method based on multi-channel integration, which can effectively solve the problems involved in the above-mentioned background technology.
[0010] To achieve the above objectives, the present invention provides the following technical solution: The first aspect of the present invention provides a recruitment information supervision platform based on multi-channel integration, comprising: a key field processing module, used to extract recruitment information metadata from various channels, extract key fields, standardize the key fields of recruitment information from various channels, generate standardized job records for each channel, and perform five-tuple modeling; a copywriting evolution generation module, used to perform cross-channel homogeneous aggregation based on the five-tuples from various channels, obtain aggregated job entities, abstract them into nodes, construct a recruitment information graph, and generate a copywriting evolution chain for the aggregated job entities; a tampering risk aggregation module, used to calculate the differences between adjacent versions on the copywriting evolution chain of the aggregated job entities, comprehensively analyze the recruitment information graph and the difference calculation results to determine the multi-dimensional risk of recruitment information, and obtain a tampering risk aggregation profile of recruitment information; and a supervision process iteration module, used to form a risk information pending group based on the tampering risk aggregation profile of recruitment information, enter manual review feedback, and adaptively iterate the recruitment information supervision process based on the manual review results.
[0011] The second aspect of this invention provides a method for supervising recruitment information based on multi-channel integration, comprising: extracting recruitment information metadata from each channel and extracting key fields; standardizing the key fields of recruitment information from each channel to generate standardized job records for each channel and performing five-tuple modeling; performing cross-channel homogeneous aggregation based on the five-tuples of each channel to obtain aggregated job entities and abstracting them into nodes to construct a recruitment information graph, generating a text evolution chain for the aggregated job entities; calculating the differences between adjacent versions on the text evolution chain of the aggregated job entities, comprehensively using the recruitment information graph and the difference calculation results to determine the multi-dimensional risk of recruitment information, and obtaining an aggregated profile of the tampering risk of recruitment information; forming a risk information pending group based on the aggregated profile of the tampering risk of recruitment information, entering manual review feedback, and adaptively iterating the recruitment information supervision process based on the manual review results.
[0012] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0013] (1) This invention provides a recruitment information supervision platform and method based on multi-channel integration. The key field processing module extracts key fields of subject, position, location, data, and time from the recruitment information metadata of various channels and performs standardization processing to generate standardized position records for each channel and perform five-tuple modeling. This eliminates aggregation errors caused by inconsistencies in field definitions across different channels at the source, ensuring that subsequent rules and models operate based on unified data definitions. The copywriting evolution generation module performs cross-channel homogeneous aggregation based on the five-tuples of each channel according to the subject ID, forming aggregated position entities and abstracting them as position nodes in the graph. At the same time, it constructs the copywriting evolution chain of the aggregated position entities, so that the release behavior of multiple channels and multiple versions is centrally expressed on a time-ordered graph structure, which facilitates the identification of complex behavior patterns such as frequent modifications and cross-channel inconsistencies. The tampering risk aggregation module calculates the difference assessment index between adjacent versions on the text evolution chain of aggregated job entities. It then combines this with subject relationships, channel relationships, and historical behavioral characteristics from the recruitment information graph to conduct multi-dimensional risk assessment, resulting in a quantifiable tampering risk aggregation profile. This allows for collaborative judgment across version, channel, and subject dimensions, improving the accuracy of identifying covert tampering and selective disclosure. The regulatory process iteration module generates a risk information pending group based on the tampering risk aggregation profile. This group undergoes manual review, collecting false alarm and confirmed samples. The module then adjusts the weights of the difference indicators, risk thresholds, and rule sensitivity in reverse, achieving adaptive iterative updates to the regulatory process. This allows the system to gradually converge its risk identification strategy during continuous operation, reducing the workload of manual review and improving the stability of the overall regulatory loop.
[0014] (2) In evaluating the real-time load quality parameters of the aggregation service gateway, this invention constructs load quality parameters by introducing multiple physical quantities such as CPU power consumption fluctuation rate, memory utilization rate, disk queue physical depth, and network card port packet loss rate. This unifies the originally scattered hardware operating states into a single control index that can directly drive the switching of aggregation modes, enabling the system to adaptively adjust between fine mode, balanced mode, and fast mode according to the load level. In this way, the aggregation service gateway can automatically reduce the computational complexity of a single aggregation during peak periods to ensure overall availability, while resuming high-precision aggregation and map updates under low load, thereby achieving a dynamic balance between computing power utilization efficiency and risk identification accuracy.
[0015] (3) Regarding the aggregation and profiling of recruitment information tampering risks, this application integrates the physical parameters of the graph, such as adjacent version difference assessment indicators, version density, version rollback rate, and cross-channel inconsistency, into a single risk profile space. This ensures that tampering risks are no longer triggered by a single field threshold, but are jointly determined by the magnitude of version changes, frequency of changes, behavioral patterns, and differences between channels. This aggregation profile outputs a continuous risk score for each aggregated job entity, facilitating regulators to quickly understand the source of risks and supporting the rule engine in developing differentiated handling strategies for different types of tampering patterns.
[0016] (4) Compared with the single-channel item-by-item review or rule judgment based on a small number of static fields commonly found in existing multi-channel supervision technologies, this application achieves improved cross-channel consistency inspection capabilities and dynamic parameter adjustment capabilities without significantly increasing collection costs by aggregating job entities through homogeneous aggregation, graph-level text evolution chain, and risk profiling mechanism with multiple parameters reuse. This enhances the system's overall ability to capture complex tampering behaviors, hidden differences across channels, and long-term behavioral patterns. Attached Figure Description
[0017] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the system module connections of the present invention.
[0019] Figure 2 This is a schematic diagram of the method steps of the present invention.
[0020] Figure 3 This is a flowchart of the preprocessing process for recruitment information.
[0021] Figure 4 A diagram illustrating the optimization of recruitment information supervision. Detailed Implementation
[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0023] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0024] Reference Figure 1 As shown, the first aspect of the present invention provides a recruitment information supervision platform based on multi-channel integration, including: a key field processing module, a copywriting evolution and generation module, a tampering risk aggregation module, and a supervision process iteration module.
[0025] The key field processing module is connected to the text evolution and generation module, the text evolution and generation module is connected to the tampering risk aggregation module, and the tampering risk aggregation module is connected to the regulatory process iteration module.
[0026] The key field processing module is used to extract the metadata of recruitment information from various channels, extract key fields, standardize the key fields of recruitment information from various channels, generate standardized job records for each channel, and perform quintuple modeling.
[0027] Based on the recruitment information metadata from various channels, key fields are extracted from the metadata fields to obtain the key fields of recruitment information from each channel.
[0028] In this embodiment, metadata fields refer to the basic attributes and structured information describing the content of recruitment information. These fields typically do not directly display specific recruitment content but rather describe the characteristics of the recruitment information itself. In recruitment information across various channels, metadata fields include, but are not limited to, the platform on which the recruitment information is published (channel ID), publication time, job type, recruiting entity (company or agency) identifier, job number, salary range, work location, job status (published, delisted, modified, etc.), document version number, and review status. These metadata fields provide a basic structural framework for the recruitment information, and their extraction can provide strong support for subsequent data processing, standardization, aggregation, and risk assessment.
[0029] Among them, key fields include at least the subject identification field and the job identification field used to identify the employer and the job entity, as well as the rights and interests substantive field used to represent salary and benefits, employment nature, fee terms, and work location. They also include the time field, channel field, and snapshot identification field used to record the evolution of document versions and cross-channel behavior trajectories. Other fields that are included in the feature set of preset regulatory rules or risk identification models can also be marked as key fields as needed.
[0030] When standardizing key fields, this regulatory platform first extracts key fields such as entity name, unified social credit code, job title, salary range, work location, employment type, fee terms, and posting time from raw recruitment information crawled from various channels based on preset field parsing rules. Redundant symbols, channel watermarks, full-width spaces, and duplicate punctuation are cleaned, and synonyms are merged. Subsequently, salary-related fields are uniformly converted into structured forms specifying upper and lower limits, currency, and pay period. The work location is mapped to a standard regional code using an administrative division database. The job title, industry category, and job category are then standardized. Text fields are mapped to a unified classification code according to a preset code table, and a unique entity identifier is searched or generated in the entity database based on the enterprise name and unified social credit code, so that the same recruitment entity can be uniformly merged under different channels; at the same time, time-related fields such as release time and delisting time are uniformly converted into timestamp representation with time zone; while outputting standardized results, the original value, intermediate parsing results and the mapping rule identifier used for each key field are retained, so that subsequent calculations can be performed directly based on a unified standard when performing cross-channel same-source aggregation, graph construction and compliance audit, and the specific source and conversion path can be traced back when needed.
[0031] Specifically, the key fields of recruitment information from various channels are standardized, and after standardization, the key fields of recruitment information from each channel meet the following constraints:
[0032] 1) Key fields of the main body of recruitment information from any channel can be mapped to a unified main body identifier after standardization.
[0033] The key fields for the entity category are primarily used to distinguish and group different employers or intermediaries. In this platform, key fields for the entity category typically include the company name, unified social credit code, organization code, business license number, and platform real-name authentication account name. By cleaning and standardizing these fields, a unique entity identifier can be generated or matched for each recruiting entity in the entity database. This allows recruitment information from the same company across different channels and formats to be uniformly linked to the same entity node, providing a stable anchor point for subsequent entity risk profiling, cross-channel behavior analysis, and blacklist / whitelist management.
[0034] 2) All location-related key fields in recruitment information from various channels are standardized and linked to a unique administrative division code or geographic grid code.
[0035] Location-related key fields refer to fields used to describe the actual work location or work area of a position, primarily used to characterize the spatial distribution features of recruitment information. In this platform, location-related key fields typically include the work city, district / county, detailed address, work area tag, and remote / on-site markers. After standardization through an administrative division database or geocoding service, these fields can be mapped to unified regional codes, latitude and longitude coordinates, or grid numbers. This allows for unified identification of the same location in various formats across different channels (e.g., Beijing Chaoyang, Beijing Chaoyang District), thereby supporting cross-channel regional distribution statistics, regional risk clustering, and refined supervision by region.
[0036] 3) All key fields of recruitment information data from various channels are represented by a standardized data structure after standardization.
[0037] Key data fields refer to those that can reflect job rights, employment conditions, or behavioral characteristics in numerical or measurable form, primarily serving as inputs for model and rule calculations. In this platform, key data fields typically include salary upper and lower limits, pay cycle, working hours, probationary period length, social security contribution ratio, fee amount, security deposit amount, and number of hires. After standardization of units, unit conversion, and interval processing, these fields can directly participate in interval comparisons, threshold determination, and statistical modeling. This helps the system identify suspicious patterns such as abnormally high salaries, excessive training fees, and excessively long working hours, thereby supporting the quantitative identification of false advertising, disguised charges, and employment risks.
[0038] 4) All time-related key fields in recruitment information from various channels are standardized and uniformly converted into a timestamp format with time zone.
[0039] Time-related key fields refer to fields used to describe key time behaviors in the lifecycle of recruitment information, mainly depicting the distribution of events such as job posting, modification, and delisting over time. In this platform, time-related key fields include posting time, modification time, delisting time, complaint time, review time, and reposting time. By uniformly converting these to timestamps with time zones or standard time formats, it is possible to sort and align text from different channels and versions on the same timeline, thereby constructing a text evolution chain that aggregates job entities. This allows for the identification of high-frequency modifications within a short period and concentrated major revisions after complaints, providing a basis for tampering risk assessment and the division of risk assessment cycles.
[0040] Furthermore, standardized job records for each channel are generated and quintuple modeling is performed. The specific analysis process is as follows:
[0041] Standardized job records include key fields for each channel, including subject-related key fields, job-related key fields, location-related key fields, data-related key fields, and time-related key fields.
[0042] Generate a unique entity ID for the current channel based on the key fields of the main body of the recruitment information from each channel, and generate a candidate ID for the current channel based on the key fields of the job position in the recruitment information from each channel.
[0043] The key fields for job categories primarily characterize the job's features in terms of business, responsibilities, and skills. In this platform, these fields typically include job title, job category, department or business line, job responsibilities, qualifications, skill requirements, work nature (full-time / part-time / internship), and employment type (formal, dispatch, outsourcing). Through synonym merging, job database mapping, and classification coding, different descriptions such as "Java Development Engineer" and "Backend Developer (Java)" can be unified under a standard job category. This allows the system to perform cross-channel aggregation, industry benchmarking analysis, and risk stratification of subdivided jobs around aggregated job entities, thereby avoiding job fragmentation and identification bias caused by differences in descriptions.
[0044] Based on the standardized job records, unique entity IDs, and job candidate IDs of each channel, a five-tuple is generated for each channel. The aggregated job entity five-tuple includes entity ID, job candidate ID, copy version ID, channel ID, and timestamp.
[0045] It should be explained that the aforementioned aggregated job entity refers to multiple job records from different channels under the same entity, but with highly similar job titles, work locations, salary ranges, and job descriptions, treated as the same job and merged into a unified job entity. This consolidates similar job information that was originally scattered across multiple channels and versions into a single job entity, facilitating unified modeling, risk control, and supervision.
[0046] The aforementioned copy version ID can be understood as an identification number for a specific version of a job posting, used to uniquely identify the copy version created when the aggregated job entity was published or modified. Within this platform, the same entity and the same aggregated job entity may publish, edit, or republish the same job posting multiple times at different times and through different channels. Each valid change (such as modifying salary, adjusting job requirements, changing work location, adding fee terms, etc.) generates a new copy version ID, which is then bound to the corresponding channel identifier, publication timestamp, standardized job record, and original page snapshot. In this way, a sequence of copy versions is formed under each aggregated job entity, linked chronologically, with each node having its own unique copy version ID. This facilitates the construction of copy evolution chains in the graph, the calculation of differences between adjacent versions, and the identification of tampering, while also enabling precise tracing of which version of content was actually published during compliance audits and evidence collection.
[0047] The copywriting evolution generation module is used to perform cross-channel homogeneous aggregation based on the five-tuples of various channels, obtain aggregated job entities and abstract them as nodes to construct a recruitment information graph, and generate copywriting evolution chains for aggregated job entities.
[0048] Specifically, cross-channel homogeneous aggregation is performed based on the five-tuples from each channel. The specific analysis process is as follows:
[0049] Extract the load status awareness dataset of the aggregation service gateway belonging to the recruitment information supervision platform, evaluate the real-time load quality parameters of the aggregation service gateway, and classify the current cross-channel same-source aggregation mode based on the real-time load quality parameters of the aggregation service gateway. The aggregation mode of the aggregated job entity includes fine mode, balanced mode and fast mode.
[0050] The specific process of dividing the above aggregation mode is as follows: the real-time load quality parameters of the aggregation service gateway are matched with the aggregation modes corresponding to the predefined load quality parameter intervals in the recruitment information supervision platform to determine the interval to which the real-time load quality parameters of the aggregation service gateway belong, and the aggregation mode corresponding to the interval is obtained.
[0051] The load quality parameter ranges include a first range, a second range, and a third range. The first range corresponds to the fine-grained mode, the second range to the balanced mode, and the third range to the fast mode.
[0052] It should be explained that the load quality parameter intervals monotonically increase according to the threshold: Assume that y1 belongs to the first interval; y2 belongs to the second interval; y3 belongs to the third interval. For any y1, y2, y3, y1 <y2<y3。
[0053] Based on the aggregation model, the five-tuples of each channel are aggregated across channels. Under the same subject ID of each channel, the job candidate ID, copy version ID, channel ID and timestamp of each channel are aggregated and aggregated into one job entity, which is called the aggregated job entity.
[0054] The aforementioned same-source aggregation process, specifically within the current aggregation window, involves the following steps: First, the regulatory platform initially groups the five-tuples from different channels according to the subject ID. Under the same subject ID, it collects all candidate IDs for each position across all channels, along with their corresponding copy version IDs, channel IDs, and timestamps. Subsequently, based on real-time load quality parameters, it selects one of three modes—fine-grained, balanced, or fast—to execute the same-source aggregation.
[0055] In fine-grained mode, the regulatory platform performs fine-grained clustering of candidate job IDs under the same subject ID based on standardized job titles, job categories, work locations, salary ranges, and text similarity. Job candidates with highly similar semantics and compatible key field values are considered as source jobs. In balanced mode, the weight of some non-key fields is weakened, and medium-grained aggregation is mainly performed using job categories, standardized job titles, and location codes. In fast mode, job candidates from different channels are quickly merged into a source set only if the subject ID is consistent and the standardized job title and location code match perfectly. For each set of job candidates obtained through source aggregation, the regulatory platform generates a unique job entity identifier as a job entity node. The text version ID, channel ID, and timestamp corresponding to all five tuples in the set are uniformly pointed to this job entity node. The resulting job entity is recorded as the aggregated job entity, which is used to carry the text evolution relationship and risk assessment results of this job under multiple channels, multiple versions, and multiple time slices in the subsequent graph.
[0056] Furthermore, the real-time load quality parameters of the aggregation service gateway are evaluated. The specific evaluation process is as follows:
[0057] The load status awareness dataset for the aggregation service gateway includes the real-time CPU power consumption fluctuation rate, real-time memory usage rate, real-time disk queue physical depth, and real-time network interface card (NIC) port packet loss rate of the aggregation service gateway.
[0058] The load status awareness dataset can be extracted from the execution records of the aggregation service gateway. Disk queue physical depth refers to the number of read / write requests waiting for the disk controller to process at a given moment; it can be understood as the length of the I / O requests queued in front of the disk. A larger queue physical depth indicates more unfinished requests on the disk, generally resulting in greater response latency and storage-side pressure.
[0059] The current CPU power consumption value of the aggregation service gateway (which can be extracted from the execution record of the aggregation service gateway) is compared with the CPU power consumption value of the aggregation service gateway at the previous moment to obtain and record the real-time CPU power consumption fluctuation rate of the aggregation service gateway.
[0060] The load status awareness dataset of the aggregation service gateway is preprocessed, including normalization and de-unitization of the aggregated job entity data. Based on the preprocessing results, weighting factors are introduced for correlation and merging to obtain the real-time load quality parameters of the aggregation service gateway. The specific analysis process is as follows:
[0061]
[0062] In the formula, Q(t) is the load quality parameter of the aggregation service gateway at time t, t is the time variable, P(t) is the CPU power consumption fluctuation rate of the aggregation service gateway at time t, M(t) is the memory utilization rate of the aggregation service gateway at time t, D(t) is the physical depth of the disk queue of the aggregation service gateway at time t, L(t) is the packet loss rate of the network card port of the aggregation service gateway at time t, w1 is the weight factor corresponding to the CPU power consumption fluctuation rate predefined in the recruitment information supervision platform, w2 is the weight factor corresponding to the memory utilization rate predefined in the recruitment information supervision platform, w3 is the weight factor corresponding to the physical depth of the disk queue predefined in the recruitment information supervision platform, and w4 is the weight factor corresponding to the packet loss rate of the network card port predefined in the recruitment information supervision platform.
[0063] In this embodiment, multivariate analysis specifically considers the correlation between these parameters. Increased CPU power consumption fluctuation often occurs in conjunction with changes in memory utilization and disk queue depth. For example, when a large number of source aggregation or graph update tasks are triggered, it will not only push the CPU into a high-frequency computing state, but also increase the memory intermediate result cache and backend write request queuing, thereby simultaneously increasing the perceived load in three dimensions and greatly reducing load quality. When the physical depth of the disk queue is large, some I / O wait will be reported to the upper layer as thread blocking or delayed aggregation result return, causing periodic busy-idle switching on the CPU side, which is reflected in a further amplification of power consumption fluctuation. An increase in network card port packet loss rate is likely to induce retransmission and connection retries, which increases the total number of requests processed per unit time, thereby putting secondary pressure on the CPU, memory, and disk queue, forming a linkage effect of abnormally amplifying the overall load on the link side, which also has a negative impact on load quality. By introducing these four types of parameters into the load quality parameters and considering their coordinated changes over time, we can more accurately distinguish between short-term single-point fluctuations and continuous systemic high-load states, avoiding misjudgments or over-adjustments caused by relying on only a single indicator.
[0064] Specifically, the recruitment information graph is constructed by aggregating job positions into entities and abstracting them into nodes. A copywriting evolution chain is then generated for each aggregated job position entity. The specific analysis process is as follows:
[0065] The recruitment information graph is constructed based on graph nodes and edge execution. The aggregated job entity graph nodes include main nodes, aggregated job entity nodes, copy version nodes, and channel nodes. The aggregated job entity edges include main-job edges, job-copy version edges, and copy version-channel edges.
[0066] The main node and the aggregated job entity node are connected via a main-job edge. The aggregated job entity node and the copywriting version node are connected via a job-copywriting version edge. The copywriting version node and the channel node are connected via a copywriting version-channel edge. It also includes a copywriting version-timestamp edge, where the copywriting version node and the timestamp are connected via a copywriting version-timestamp edge. In the recruitment infographic, the copywriting version-timestamp edge represents the evolution chain of copywriting versions. This edge allows tracking the historical changes of the same job position, clarifying when each copywriting version was released, modified, or removed. Therefore, time here is not only part of the data but also a timeline describing the change process, determining the order and changes between versions.
[0067] Arrange the copywriting version nodes under the aggregated job entity node in order of copywriting version timestamp to form the copywriting evolution chain of the aggregated job entity node.
[0068] Furthermore, the aggregation of job postings into aggregated entities and their abstraction into nodes is used to construct a recruitment information graph. This process generates a copywriting evolution chain for the aggregated job posting entities and also includes:
[0069] When the aggregation mode is displayed as fine-grained mode or balanced mode, the aggregation service gateway hashes and shards the main node, aggregation position entity node and copy version node according to the main ID, and stores them on each storage node of the aggregation service gateway.
[0070] When the aggregation mode is displayed as fast mode, the read and write frequencies of each storage node of the aggregation service gateway are extracted. Based on the read and write frequencies of each storage node, hotspot subjects / aggregation job entities are identified, and the storage nodes to which the hotspot subjects / aggregation job entities belong are arranged from largest to smallest to form a storage node load sequence. Based on the storage node load sequence, the migration strategy of hotspot subjects / aggregation job entities is triggered.
[0071] The read / write frequency of each storage node is compared with the predefined read / write threshold frequency. The main / aggregated job entities under several storage nodes whose read / write frequency is greater than the read / write threshold frequency are recorded as hotspot main / aggregated job entities, thus completing the identification of hotspot main / aggregated job entities.
[0072] After identifying the hot topics / aggregated job entities, the aggregation service gateway selects several target storage nodes based on the storage node load sequence. These target nodes have read / write frequencies lower than the preset read / write threshold frequency and their storage capacity and network bandwidth are within safe margins. A migration mapping table for the hot topics / aggregated job entities is then constructed. Subsequently, short-term write rate limiting or locking control is applied to the hot topics / aggregated job entities to be migrated, as well as their associated job entity nodes, document version nodes, and relationship edges. Without affecting read requests, the data of these nodes is migrated to the target storage nodes in batch replication or incremental synchronization. After the migration is completed, the total number of nodes, total number of edges, and checksums on the source and target nodes are compared to ensure the consistency and integrity of the migrated data.
[0073] After the data consistency verification passes, the routing table or shard metadata is updated, and subsequent read and write requests for hot subjects / aggregated roles are directed to the new target storage node. A read-only copy of the old storage node is retained for a preset buffer period to meet rollback requirements. Finally, the load index and read / write frequency of the source storage node that has completed the migration are recalculated so that the load distribution of each storage node can be dynamically corrected in the next round of hot spot identification and migration cycle, so as to achieve continuous balanced scheduling of hot subjects / aggregated roles within the cluster.
[0074] In summary, the overall process of aggregating job postings and processing and aggregating recruitment information is as follows: Figure 3 As shown, Figure 3 The flowchart for the preprocessing of recruitment information is as follows: First, the key field processing module generates standardized job records and constructs a five-tuple model. Next, through cross-channel homogeneous aggregation, different aggregation modes are selected based on system load: a fine-grained mode performs high-precision aggregation based on similarity and standardized job titles; a balanced mode performs medium-precision aggregation based on job category and location coding; and a fast mode performs rapid aggregation based on direct matching of the subject ID and job title. The aggregated unified job entity is then bound to each document version, ultimately constructing a complete recruitment information graph to ensure information consistency and traceability, supporting subsequent risk assessment and analysis.
[0075] The tampering risk aggregation module is used to calculate the differences between adjacent versions on the text evolution chain of aggregated job entities, and to make multi-dimensional risk judgments on recruitment information by combining the recruitment information graph and the difference calculation results, so as to obtain a tampering risk aggregation profile of recruitment information.
[0076] Specifically, a multi-dimensional risk assessment of recruitment information is conducted based on the comprehensive recruitment information graph and the difference calculation results, resulting in an aggregated profile of the risk of tampering with recruitment information. The specific analysis process is as follows:
[0077] On the text evolution chain of the aggregated job entity, the characteristic parameters of each adjacent version of the aggregated job entity are extracted, and the difference evaluation index of each adjacent version of the aggregated job entity is calculated.
[0078] The aforementioned feature parameters represent the text similarity differences between adjacent versions of the aggregated job posting entity, the release time interval between adjacent versions of the aggregated job posting entity, the version modification frequency between adjacent versions of the aggregated job posting entity, and the cumulative amount of changes between adjacent versions of the aggregated job posting entity. The text similarity difference is calculated by determining the text similarity between two adjacent versions, for example, using cosine similarity to quantify the degree of change in the text content. The release time interval, version modification frequency, and cumulative amount of changes can be extracted from the public records of the recruitment information supervision platform. The cumulative amount of changes refers to the cumulative changes in all key fields (such as salary, job description, work location, etc.) between two adjacent versions of the recruitment information text. Specifically, it measures the comprehensive changes in the values, content, or format of all fields between the two versions.
[0079] During the risk assessment period, the version density of aggregated job entities, the version rollback rate of aggregated job entities, and the cross-channel inconsistency of recruitment information are extracted based on the recruitment information graph representation.
[0080] The risk assessment cycle refers to the time interval at which the regulatory platform centrally collects, updates, and scores various risk indicators for aggregated job entities and their related entities according to a preset time window (such as every 5 minutes, every hour, or every day). By setting a risk assessment cycle, a balance can be struck between not consuming resources too frequently and reflecting the latest tampering and abnormal behavior in a timely manner, so that the risk score can be iteratively updated after each cycle.
[0081] The above version rollback rate is expressed as follows: for the same evolutionary chain, if a certain version v k+2 Compared to earlier versions v k Highly similar, and similar to the intermediate version v k+1 Significant differences are recorded as a single rollback event, where 'k' represents the version index, indicating the current version. The version rollback rate is defined within the time window.
[0082] ρ rollback =N rollback / N ver ;
[0083] Where, ρ rollback Expressed as version rollback rate, N rollback This is expressed as the number of rollback events detected within a selected time window. Whenever version v occurs... k+1 and v k The differences are significant, while v k+2 And vk A highly similar pattern of changing and then reverting is counted as one rollback event, and these are accumulated to obtain N. rollback N ver It is expressed as the total number of copywriting versions (or the total number of version changes) that appear on this evolutionary chain within the same time window.
[0084] The aforementioned cross-channel inconsistency is expressed as the set of versions from each channel within the same time slice (or approximately a time window) {v c Define cross-channel inconsistency: Where Δch represents the cross-channel inconsistency, C is the number of channels for this position in the current time slice, and 2 / C(C-1) is a normalization of the summation result to obtain an average difference value. The total differences are summed and then divided by the number of all channel pairs to ensure that the cross-channel inconsistency is averaged, not cumulative. d(·,·) is the difference based on key fields (which can be a combination of previous field differences and text similarity), and c... x c y This indicates different channel pairs. When calculating cross-channel differences, two different channels (x and y represent the channel pair numbers) are selected, and the differences between their copy versions are compared.
[0085] Example:
[0086] A certain aggregation position published copy on three channels during the current time slot:
[0087] Channel 1: c1, Channel 2: c2, Channel 3: c3; the three corresponding copywriting versions are: v c1 v c2 v c3 .
[0088] First, calculate the difference d(·,·) between any two channels' copywriting, for example (the numbers are just examples):
[0089] The degree of difference between channel 1 and channel 2: d(v c1 ,v c2 = 0.2.
[0090] Difference between Channel 1 and Channel 3: d(v c1 ,v c3 = 0.5.
[0091] Difference between Channel 2 and Channel 3: d(v c2 ,v c3 = 0.3.
[0092] At this time:
[0093] The total number of channels C = 3. There are three pairs of channels that satisfy x < y: (c1, c2), (c1, c3), (c2, c3). First, sum up the difference degrees of the three pairs:
[0094]
[0095] Then multiply by the previous coefficient:
[0096]
[0097] Finally, obtain the cross-channel inconsistency degree: Δch = 1 / 3 × 1.0 ≈ 0.333.
[0098] Perform data preprocessing on the difference evaluation index of each adjacent version of the aggregated job entity, the version density of the aggregated job entity, the version rollback rate of the aggregated job entity, and the cross-channel inconsistency degree of the recruitment information. The data preprocessing of the aggregated job entity includes normalization processing and unit removal processing.
[0099] Introduce a weight factor based on the data preprocessing results for correlation and merging to obtain the tampering risk aggregation portrait of the recruitment information. The specific implementation process is as follows:
[0100]
[0101]
[0102] In the formula, Rj is the tampering risk aggregation portrait of the recruitment information, Dd(k) is the difference evaluation index between the k-th and the (k + 1)-th versions of the aggregated job entity, k is the version number, k = 1, 2, 3,..., B, B is the total number of versions, λv is the version density of the aggregated job entity, ρ rollbackLet Δch be the version rollback rate of the aggregated job entity, Δch be the cross-channel inconsistency of the recruitment information, S(k) be the text similarity difference between the k-th and k+1-th versions of the aggregated job entity, ΔT(k) be the release time interval between the k-th and k+1-th versions of the aggregated job entity, F(k) be the version modification frequency between the k-th and k+1-th versions of the aggregated job entity, C(k) be the cumulative change between the k-th and k+1-th versions of the aggregated job entity, y1 be the weight factor corresponding to the predefined average difference assessment index in the recruitment information supervision platform, and y2 be the weight factor corresponding to the average difference assessment index in the recruitment information supervision platform. The predefined weight factors are: version density, version rollback rate, cross-channel inconsistency, text similarity difference, release time interval, version modification frequency, and cumulative change amount.
[0103] In this embodiment, multivariate analysis specifically considers the correlation between these parameters. The adjacent version difference assessment index and version density often appear in tandem. When a certain aggregated position entity frequently releases new versions in a short period of time, it may increase the version density per unit time and easily accumulate multiple large-scale content shifts, making the behavior of making many and drastic changes amplified in the aggregated profile. When the version rollback rate increases, it usually means that there is a back-and-forth pattern of normal description, suspicious description, and close to the original description in multiple versions. This pattern relies on the adjacent version difference assessment index to identify suspicious jumps on the one hand, and on the other hand, it also relies on the version density to form a relatively concentrated trial-and-retraction segment on the time axis. There is also a superposition effect between cross-channel inconsistency and adjacent version difference assessment index. When the same position makes frequent and large-scale modifications in the time dimension and shows obvious inconsistencies in the channel dimension, the aggregated profile no longer regards it as a normal modification in a single channel, but incorporates the differences of multiple time slices and multiple channels into the assessment, so that the tampering risk forms a convergence direction in the perspectives of the subject, position, and channel, thereby enhancing the sensitivity and discrimination of the aggregated profile to complex tampering behavior.
[0104] The regulatory process iteration module is used to aggregate the risk profile of tampering in recruitment information to form a risk information pending group, which then enters the manual review feedback. Based on the results of the manual review, the recruitment information regulatory process is adaptively iterated.
[0105] Furthermore, the recruitment information supervision process is adaptively iterated based on the results of manual review. The specific analysis process is as follows:
[0106] The risk profile of tampering with recruitment information is aggregated and compared with the predefined first and second thresholds for tampering risk. Based on the comparison results, a risk information pending group is formed and uploaded to the recruitment information supervision platform for manual review and feedback.
[0107] The comparison results above include a first comparison result, a second comparison result, and a third comparison result. The first comparison result shows that the aggregated profile of the recruitment information tampering risk is greater than the first threshold for tampering risk. The second comparison result shows that the aggregated profile of the recruitment information tampering risk is less than or equal to the first threshold for tampering risk, and greater than the second threshold for tampering risk. The third comparison result indicates that the aggregated profile of the recruitment information tampering risk is less than or equal to the second threshold for tampering risk.
[0108] The job posting is marked as "violation" when the comparison result is the first result, "abnormal" when the comparison result is the second result, and "normal" when the comparison result is the third result. Therefore, the "risk information pending" group includes job postings marked as "violation," "abnormal," and "normal."
[0109] The results of manual review include confirmed violations of risk information, general anomalies of risk information, and false alarms of risk information.
[0110] The set of recruitment information that is displayed as a false risk information by manual review is called the false alarm set. The recruitment information supervision process is adaptively iterated based on the display representation of the false alarm set.
[0111] In one example implementation, the regulatory platform first collects recruitment information that has been manually reviewed and confirmed as false alarms into a false alarm set. It then extracts display characteristics such as difference assessment indicators, version density, version rollback rate, cross-channel inconsistency, and high-risk field insertion rate of the aggregated job entities in the false alarm set, and statistically analyzes their distribution characteristics and threshold triggering situations in each dimension. Subsequently, it compares and analyzes the feature distribution of the false alarm set with the feature weights and alarm thresholds used in the current risk identification model and rule engine to identify which feature combinations, weight settings, or threshold ranges are more likely to cause false alarms for normal recruitment information, and marks these combinations as biased sensitive configurations.
[0112] Based on this, the regulatory platform lowers the weight of relevant features, moderately raises overly aggressive alarm thresholds, or adds whitelist feature fragments and false alarm pattern filtering conditions to the front end of the rules to reduce repeated triggering of such patterns. After completing the parameter update, the regulatory platform identifies newly added recruitment information with the new configuration in the next risk assessment cycle and continuously monitors new manual review results. The incremental increase of new false alarm samples is incorporated into the false alarm set, forming a closed-loop iterative process of false alarm set update, feature representation correction, threshold and weight adaptive adjustment, and a new round of identification and review. This allows the recruitment information supervision process to gradually reduce the false alarm rate and improve the targeting and stability of risk identification in multiple iterations.
[0113] It should be explained that, in an example embodiment, the above-mentioned adjustment range constraint is manifested in the fact that after the platform includes the false alarm set into the training library, it first calculates the trigger ratio p of each feature f in the false alarm samples. fp (f) The trigger ratio p in the normal samples that passed the audit. norm (f), when the difference between the two is p fp (f)-p norm (f) When the weight of the feature is greater than the preset bias threshold Δp, the weight of the feature is adjusted accordingly. The threshold is lowered, with the adjustment coefficient η limited to the range of [0.2, 0.5], and an upper limit is set for the weight reduction in a single round, for example, not exceeding 10%–30% of the original value, thus making the degree of reduction quantifiable and controllable; for the overall alarm threshold T risk The platform calculates the false alarm rate r based on the actual false alarm rate during the current assessment period. fp With the target false alarm rate r fp * The difference is calculated according to Adjustments are made upwards or downwards, with the coefficient h controlling the single-round threshold adjustment to not exceed 3%–10% of the original threshold, avoiding excessive correction at once. After the above quantitative adjustment, the updated feature weights and alarm thresholds take effect in the next risk assessment cycle, and the above calculation is iteratively repeated in combination with the new false alarm set statistical results, so that the degree of upward / downward adjustment is constrained by explicit formulas and percentage upper limits.
[0114] Upon confirmation of a violation in risk information, the platform will immediately trigger a high-risk warning based on the results of manual review, marking the information as non-compliant and prohibiting its continued posting on the recruitment platform. The monitoring system will automatically generate a detailed violation report, listing the specific manifestations of the violation, such as false information, illegal fee clauses, or non-compliant job descriptions, for relevant regulatory departments to handle further. Simultaneously, the job posting will be blacklisted to prevent similar violations by the same entity or position from recurring. This information will be incorporated into the risk monitoring system to further strengthen the monitoring of the entity and prevent the spread of potential compliance risks.
[0115] In cases of generally unusual risk information, the platform will issue a warning, alerting relevant personnel that the position may contain non-standardized behaviors, such as inconsistent salary descriptions, vague job requirements, or slightly misleading information. The monitoring system will flag the information based on manual review results and send a warning to the recruiter, requiring them to modify or provide further clarification within a specified timeframe. Simultaneously, the system will record this unusual information as monitoring data for subsequent analysis and decision-making, ensuring rapid identification and timely intervention of similar anomalies, thereby reducing potential risks in recruitment information.
[0116] In summary, the process of aggregating and iteratively optimizing the risk of tampering with recruitment information is as follows: Figure 4 As shown Figure 4 The diagram illustrates the optimization of recruitment information supervision. First, the tampering risk aggregation module calculates the differences between adjacent versions and combines this with a comprehensive recruitment information graph to perform multi-dimensional risk assessment, generating a tampering risk aggregation profile for the recruitment information. Next, the system enters a manual review stage, determining whether there are false alarms based on the review results. If there are false alarms, the samples are added to the false alarm set, affecting subsequent iterations; if there are no false alarms, a recruitment information warning is triggered. Feedback from the false alarm set is used to adjust risk thresholds, feature weights, and rule sensitivity, thereby enabling adaptive updates and optimization of the system, ensuring that the accuracy and effectiveness of risk identification gradually improves during operation.
[0117] Reference Figure 2 As shown, the second aspect of the present invention provides a method for supervising recruitment information based on multi-channel integration, including: extracting recruitment information metadata from each channel, extracting key fields, standardizing the key fields of recruitment information from each channel, generating standardized job records for each channel, and performing quintuple modeling.
[0118] Based on the five-tuples of each channel, cross-channel homogeneous aggregation is performed, which is aggregated into aggregated job entities and abstracted into nodes to construct a recruitment information graph, and a copywriting evolution chain is generated for the aggregated job entities.
[0119] The differences between adjacent versions are calculated on the text evolution chain of the aggregated job entities. The recruitment information graph and the difference calculation results are used to make a multi-dimensional risk assessment of the recruitment information and obtain an aggregated profile of the recruitment information tampering risk.
[0120] Based on the risk profile of tampering with recruitment information, a risk information pending group is formed, which then enters the manual review and feedback process. The recruitment information supervision process is adaptively iterated based on the results of the manual review.
[0121] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined by the present invention, they should all fall within the protection scope of the present invention.
Claims
1. A recruitment information supervision platform based on multi-channel integration, characterized in that: include: The key field processing module is used to extract the metadata of recruitment information from various channels, extract key fields, standardize the key fields of recruitment information from various channels, generate standardized job records for each channel, and perform five-tuple modeling. The copywriting evolution generation module is used to perform cross-channel homogeneous aggregation based on the five-tuples of various channels, obtain aggregated job entities and abstract them as nodes to construct a recruitment information graph, and generate copywriting evolution chains for aggregated job entities; The tampering risk aggregation module is used to calculate the differences between adjacent versions on the text evolution chain of aggregated job entities, and to make multi-dimensional risk judgments on recruitment information by combining the recruitment information graph and the difference calculation results, so as to obtain a tampering risk aggregation profile of recruitment information. The regulatory process iteration module is used to aggregate risk profiles based on the tampering of recruitment information to form a risk information pending group, which then enters the manual review feedback. Based on the results of the manual review, the recruitment information regulatory process is adaptively iterated.
2. The recruitment information supervision platform based on multi-channel integration as described in claim 1, characterized in that: The key fields of recruitment information from each channel are standardized, and the standardized key fields of recruitment information from each channel meet the following constraints: 1) Key fields of the main body of recruitment information from any channel can be mapped to a unified main body identifier after standardization processing; 2) After standardization, all location-related key fields in any recruitment information from various channels are associated with a unique administrative division code or geographic grid code; 3) All key fields of recruitment information data from various channels are represented by a standardized data structure after standardization; 4) All time-related key fields in recruitment information from various channels are standardized and uniformly converted into a timestamp format with time zone.
3. The recruitment information supervision platform based on multi-channel integration as described in claim 2, characterized in that: The process of generating standardized job records for each channel and performing quintuple modeling is as follows: The standardized job records include key fields for the main body, job, location, data, and time for each channel. Generate a unique entity ID for the current channel based on the key fields of the main body of the recruitment information from each channel, and generate a candidate ID for the current channel based on the key fields of the job position in the recruitment information from each channel. Based on the standardized job records, unique entity IDs, and job candidate IDs of each channel, a quintuple is generated for each channel. The quintuple includes entity ID, job candidate ID, copy version ID, channel ID, and timestamp.
4. The recruitment information supervision platform based on multi-channel integration as described in claim 1, characterized in that: The specific analysis process for cross-channel homogeneous aggregation based on the five-tuples of each channel is as follows: Extract the load status awareness dataset of the aggregation service gateway to which the recruitment information supervision platform belongs, evaluate the real-time load quality parameters of the aggregation service gateway, and classify the current cross-channel same-source aggregation into aggregation modes based on the real-time load quality parameters of the aggregation service gateway. The aggregation modes include fine mode, balanced mode and fast mode. Based on the aggregation model, the five-tuples of each channel are aggregated across channels. Under the same subject ID of each channel, the job candidate ID, copy version ID, channel ID and timestamp of each channel are aggregated and aggregated into one job entity, which is called the aggregated job entity.
5. The recruitment information supervision platform based on multi-channel integration as described in claim 4, characterized in that: The evaluation process for the real-time load quality parameters of the aggregation service gateway is as follows: The load status awareness dataset of the aggregation service gateway includes the real-time CPU power consumption fluctuation rate, the real-time memory usage rate, the real-time disk queue physical depth, and the real-time network interface card port packet loss rate of the aggregation service gateway. The load status awareness dataset of the aggregation service gateway is preprocessed, including normalization and de-normalization. Based on the data preprocessing results, a weight factor is introduced for correlation and merging to obtain the real-time load quality parameters of the aggregation service gateway.
6. The recruitment information supervision platform based on multi-channel integration as described in claim 1, characterized in that: The process of obtaining aggregated job entities and abstracting them into nodes to construct a recruitment information graph, and generating a copywriting evolution chain for the aggregated job entities, is as follows: The recruitment information graph is constructed based on graph nodes and edge execution. The graph nodes include main nodes, aggregated job entity nodes, copy version nodes, and channel nodes. The edges include main-job edges, job-copy version edges, and copy version-channel edges. The main node and the aggregated job entity node are connected through the main-job edge, the aggregated job entity node and the copy version node are connected through the job-copy version edge, and the copy version node and the channel node are connected through the copy version-channel edge. Arrange the copywriting version nodes under the aggregated job entity node in order of copywriting version timestamp to form the copywriting evolution chain of the aggregated job entity node.
7. The recruitment information supervision platform based on multi-channel integration as described in claim 6, characterized in that: The process of obtaining aggregated job entities and abstracting them into nodes to construct a recruitment information graph, and generating a copywriting evolution chain for the aggregated job entities, also includes: When the aggregation mode is displayed as fine-grained mode or balanced mode, the aggregation service gateway hashes and shards the main node, aggregation position entity node and copy version node according to the main ID, and stores them on each storage node of the aggregation service gateway. When the aggregation mode is displayed as fast mode, the read and write frequencies of each storage node of the aggregation service gateway are extracted. Based on the read and write frequencies of each storage node, hotspot subjects / aggregation job entities are identified, and the storage nodes to which the hotspot subjects / aggregation job entities belong are arranged from largest to smallest to form a storage node load sequence. Based on the storage node load sequence, the migration strategy of hotspot subjects / aggregation job entities is triggered.
8. The recruitment information supervision platform based on multi-channel integration as described in claim 1, characterized in that: The comprehensive recruitment information map and the difference calculation results are used to determine the multi-dimensional risks of recruitment information, resulting in an aggregated profile of recruitment information tampering risks. The specific analysis process is as follows: On the copywriting evolution chain of the aggregated job entity, extract the characteristic parameter representation of each adjacent version of the aggregated job entity, and calculate the difference evaluation index of each adjacent version of the aggregated job entity. During the risk assessment period, the version density of aggregated job entities, the version rollback rate of aggregated job entities, and the cross-channel inconsistency of recruitment information are extracted based on the recruitment information graph representation. The data preprocessing includes normalization and de-unitization. The evaluation indicators of the differences between adjacent versions of the aggregated job entity, the version density of the aggregated job entity, the version rollback rate of the aggregated job entity, and the cross-channel inconsistency of recruitment information are evaluated. Based on the data preprocessing results, weighting factors are introduced for correlation and merging to obtain an aggregated profile of the risk of tampering with recruitment information.
9. The recruitment information supervision platform based on multi-channel integration as described in claim 1, characterized in that: The adaptive iterative analysis process for the recruitment information supervision process based on manual review results is as follows: The risk profile of tampering with recruitment information is aggregated and compared with the predefined first and second thresholds for tampering risk. Based on the comparison results, a risk information pending group is formed and uploaded to the recruitment information supervision platform for manual review and feedback. Manual review results include confirmed violations of risk information, general anomalies in risk information, and false alarms in risk information; The set of recruitment information that is displayed as a false risk information by manual review is called the false alarm set. The recruitment information supervision process is adaptively iterated based on the display representation of the false alarm set.
10. A recruitment information supervision method based on multi-channel integration, applied to the recruitment information supervision platform based on multi-channel integration as described in any one of claims 1-9, characterized in that: include: Extract key fields from recruitment information metadata from various channels, standardize the key fields of recruitment information from various channels, generate standardized job records for each channel, and perform quintuple modeling. Based on the five-tuples of each channel, cross-channel homogeneous aggregation is performed to obtain aggregated job entities and abstract them as nodes to construct a recruitment information graph, and generate a copywriting evolution chain for the aggregated job entities. The differences between adjacent versions are calculated on the text evolution chain of the aggregated job entity. The recruitment information graph and the difference calculation results are combined to make a multi-dimensional risk assessment of recruitment information and obtain an aggregated profile of the risk of tampering with recruitment information. Based on the risk profile of tampering with recruitment information, a risk information pending group is formed, which then enters the manual review and feedback process. The recruitment information supervision process is adaptively iterated based on the results of the manual review.
Citation Information
Patent Citations
Mass user multi-channel data management method and electronic equipment
CN117725098A
Human resource text information management method and device based on big data and medium
CN119988580A