A machine learning-based natural language processing (NLP) system for automatically detecting gender bias in job postings
A machine learning-based NLP system addresses sex-specific biases in job postings by providing real-time feedback and suggestions for neutral language, enhancing inclusivity and diversity in recruitment processes.
Patent Information
- Application Number
- DE202025102481
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2035-05-31
AI Technical Summary
Existing automated systems for reviewing and publishing job postings often introduce sex-specific biases in location descriptions, which are difficult to detect and correct, leading to reduced inclusivity and diversity in recruitment processes.
A machine learning-based NLP system that utilizes deep learning models and linguistic theories to identify and mitigate sex-specific distortions in location indicators, providing real-time feedback and suggestions for neutral alternatives, and continuously learns from user feedback to improve accuracy.
The system effectively reduces the need for manual reviews, enhances sex-neutral location descriptions, promotes diversity and inclusion, and integrates seamlessly into existing recruiting platforms for scalable bias detection and correction.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present invention relates to natural language processing (NLP) systems and methods, and more particularly to a machine learning controlled NLP system for detecting sex-specific distortions in location indicators. More specifically, the process is automated to recognize and mitigate sex-specific distortions in location descriptions, thereby promoting sex neutrality and diversity in the adjustment process.In recent years, the widespread use of online recruiting platforms has led to the increasing use of automated systems for the review and publication of site prescriptions. These systems have improved the efficiency of the adjustment methods, but have also raised considerable concerns regarding fitness, inclusivity and pre-compliance. Among the various forms of pre-inclusion, sex-specific pre-inclusion is still a widespread problem, often arising subtilly in the language of the site indicators. Words or formulations that are perceived as sex-specific encoded may inadvertently prevent users of particular sex from advertising, thus compromising the variety and compliance with the setting.Manual checking of place assignments for sex-specific differences is not only time-consuming and subjective, but also not scalable, especially for companies that handle thousands of place descriptions. Existing tools and techniques are only limited in their ability to identify nuancing forms of pre-agreement, especially those that are attributable to complex linguistic patterns. Therefore, there is an increasing need for a smart scalable solution that can detect and mitigate sex-specific distortions in location indicators before they are published, thus ensuring speech neutrality and promoting the diversity of users.To address these challenges, the present invention discloses a machine learning based natural language processing (NLP) system (100) for automatically identifying sex-specific pre-judgments in place pronunciations. The system utilizes deep learning models trained on large, annotated sex specific language Korspora to analyze text content, mark sex specific expressions, and provide real-time feedback or suggestions for neutral alternatives. By integrating linguistic theories with advanced ML algorithms, the system provides accurate, context aware recognition and promotes fair tuning practice.An object of the present disclosure is that the system allows for automatic recognition of sex specific preurites in real-time site prescriptions, thus reducing the need for manual reviews.It is another object of the present disclosure that the system helps organizations create more sex-neutral location descriptions and thus promote diversity and inclusion in adjustment practice.Another object of the present disclosure is that the system continuously learns and improves from the feedback of the users, thereby increasing accuracy over time.Another object of the present disclosure is that personal responsible and staff agents can quickly recognize and correct pre-taken formulations, thereby accelerating the process of site writing.Another object of the present disclosure is that the system can be seamlessly integrated into existing recruiting platforms, thereby being able to scale across organizations.Another object of the present disclosure is that the system, through the use of NLP and contextual word embeddings, detects subtil distortions that could escape other traditional methods.Another object of the present disclosure is that the distortion detection and reporting module provides detailed feedback and suggestions for improving location descriptions.Another object of the present disclosure is the intuitive user interface that provides easy interaction for personnel responsible and does not require technical knowledge for effective use thereof.The present invention relates generally to a machine learning based NLP system for automatically identifying sex specific preurites in site displays. The system utilizes a combination of pre-trained NLP models, sex specific preurition algorithms, and machine learning classification methods to identify sex specific speech, stereo types, and implicit preuritions in spot descriptions.One embodiment of the present invention is the pre-processing module that processes the raw text of the location pronunciation to remove disturbing factors such as displays, formatting and irrelevant data while preserving the core content of the text.Another embodiment of the invention is that the system extracts features from the text of the sentence, including keyword-based features (e.g., male or female encoded words), sentence structure patterns, and linguistic cues that may indicate sex-specific alignment.Another embodiment of the invention is that the machine learning model trained on a marked dataset with sex-specific and non-sex-specific location descriptions classifies the location pronunciation as sex-specific or non-sex-specific based on the extracted features. The model may be continuously updated with new data to improve its recognition accuracy.Another embodiment of the invention is that once a preference is determined, the system highlights the specific areas of the location pronunciation where a sex-specific preference was determined and provides suggestions for neutral language alternatives.Another embodiment of the invention is that a user friendly interface allows personal responsible and administrators of site indicators to upload site descriptions for analysis. The system provides real-time feedback and allows for rapid changes to ensure more integrative speech.Another embodiment of the invention is that the system can be incorporated into site writing platforms and recruiting software, which allows automatic analysis and feedback as part of the site writing process.The present invention relates to a machine learning-based NLP system developed for the automatic recognition of sex-specific lectures in site prescriptions. It comprises several key modules which cooperate to identify and mitigate sex-specific preuritions. The pre-processing module corrects and normalizes the raw text while the feature extraction module identifies sex specific encoded speech, stereo types, and contextual distortions. The machine learning model classifies the location pronunciation as pre-entered or un-entered based on the extracted features. A distortion detection and reporting module highlights distorted formulations and suggests neutral alternatives, and a user interface provides an accessible platform for personnel responsible for checking and modifying spot descriptions. The integration module ensures that the system can be integrated seamlessly into existing recruiting platforms to perform real-time analyses.FIG. 1 shows that the machine learning-based NLP system for automatic recognition of sex-specific preurites in site displays is comprised of multiple key modules, each designed for a particular aspect of the process.The first module, preprocessing, prepares the raw text of the location descriptor for analysis. This module begins with tokenizing the text, i.e., it is broken down into smaller units (words or subwords) to better understand its structure. The text is then normalized by converting all characters to lower case letters, removing special characters, and correcting spell errors. To improve efficiency, frequently occurring stop words (such as "the", "and" or "is") are removed and words are reduced to their basic form by lemmasing (e.g., "run" becomes "run"). In addition, the text is divided into meaningful sentence units, and irrelevant contents such as formatting or advertising are removed. This preprocessing ensures that only the essential text remains, which makes it easier for the subsequent modules to analyze the speech effectively.The module feature extraction follows, which focuses on the identification and extraction of key features from the pre-processed text. This module identifies sex specific words -- terms typically associated with maleity or femaleity, such as "strong",... "penetrable", or "competitive"(male) and "fürsorglich",,unterstützend' or "feeling" (female). Also analyzed are sentence structures and patterns that might indicate sex-specific preuritions, e.g., the use of sex-specific pronouns or formulations such as "preferably female" or "male candidate.". The system searches for stereotyped formulations that enhance traditional sex roles, e.g., "we search for a strong male guide force" or "we search for a woman who is responsible.". The module uses contextual analysis with word embeddings such as Word2Ve, GloVe, or BERT that detect the meaning of words based on the context surrounding them. This contextual approach helps to detect subtilous distortions that might not immediately occur with isolated keywords alone. In addition, the module examines the word frequency and common occurrence patterns to recognize potential distortions, e.g., when items associated with guidance are disproportionately strongly associated with male-coded speech. The extracted features are then prepared for analysis by the machine learning model.The heart of the system is the Machines Learning Model module. This model is responsible for analyzing the features extracted from the text of the location descriptor and classifying them as either pre-accepted or un-pre-accepted.It is trained on a large, marked dataset of location indicators that contains both sex specific and neutral examples. In the training process, the extracted features, such as the presence of sex specific words, sentence structure, and context information, are fed into a machine learning algorithm. The system uses supervised learning techniques such as logistic regression, support vector machines (SVM), random forests or deep learning models such as BERT or LSTM. These models are specifically tuned to the recognition of sex specific distortions in site descriptions. After training, the model classifies new location assignments based on the detected features and provides a bias score that indicates the probability of sex-specific distortion. This module serves as the core of the decision making and determines whether a site pronunciation contains sex specific formulations.The bias detection and reporting module cooperates with the machine learning model to provide useful insights. When a distorted location pronunciation is detected, this module highlights the specific words, phrases, or phrases that contribute to the distortion. It generates a detailed report containing the exact position of the distorted speech and providing suggestions for neutral language alternatives. For example, if the site pronunciation contains the phrase "strong male candidate", the system may recommend changing it to "strong candidate" to eliminate sex-specific distortion. Based on this feedback, the staff responsible and the managers of the site indicators can review the site descriptions to ensure that they are inclusive and sex neutral before being published.Finally, the User Interface (UI) module provides a user-friendly platform for personal responsibles and personnel brokers to interact with the system. Users can upload place descriptions by either inserting the text directly into a dedicated field or uploading the text file. The system then analyzes the location pronunciation in real time and provides feedback about determined sex-specific distortions. The user interface displays an assessment of the distortions and highlights the specific areas of the location writing that require attention. Moreover, the users may identify certain terms or phrases manually as being pre-accepted, so that the system may learn from the feedback of the users and continuously improve their accuracy. The user interface is intuitive and user-friendly, so that personal responsibility can seamlessly incorporate sex-neutral formulations into place pronunciations.The system also includes an integration module that enables seamless integration of sex specific vortical recognition functions into existing recruiting platforms, Advertising Tracking Systems (ATS), or Site Writing Websites. This module ensures that recognition of sex-specific lectures becomes an automated part of the workflow in the place prescription, so companies can recognize potential lectures in real time and adjust their place prescriptions before they are published publicly. By integrating with these platforms, the system ensures that recognition of sex-specific pre-judgments is a consistent and scalable part of the adjustment process, thus contributing to the promotion of diversity and inclusivity in adjustment practice.
Claims
A machine learning controlled natural language processing (NLP) system (100) for automatically recognizing sex-specific distortions in location displays, comprising: a) a pre-processing module configured to process raw text data from a location pronunciation by performing tokenization, text normalization, stopword removal, lemmaization, sentence segmentation, and noise removal to extract meaningful linguistic content; b) a feature extraction module that identifies sex-encoded words, analyzes sentence structures and linguistic patterns, recognizes stereo-type expressions, and performs context analysis to extract features indicative of sex-specific distortions in the location pronunciation; c) a machine learning model trained on tagged data to classify the site pronunciation as sex-specific pre-entered or not pre-entered based on the extracted features, the model selected from a group consisting of logistic regression, support vector machine, random forest, and deep learning models; d) a pronouncing recognition and notification module that highlights the sex-specific language location in the site pronunciation, generates a pronouncing score, and provides proposed neutral language alternatives; e) a user interface configured to receive the text of the location prescription, display the recognized sex-specific distortions and the proposed corrections, and allow users to manually mark distorted terms to improve the accuracy of the system; and f) an integration module configured to integrate the system into existing location prescription platforms, Answer Tracking Systems (ats), or Recruiting software to recognize sex-specific distortions in real time during the location prescription process.The system (100) of claim 1, wherein the preprocessing module further comprises a submodule that removes displays, formatting code, and irrelevant content from the location descriptor prior to the analysis.The system (100) of claim 1, wherein the feature extraction module uses word embeddings selected from the group consisting of Word2Ve, GloVe, and BERT to perform contextual analysis of the extracted features.The system (100) of claim 1, wherein the machine learning model is a deep learning model, further comprising a pre-trained language model selected from the group consisting of BERT, GPT, and LSTM, fine tuned for recognition of sex-specific distortions in location descriptions.The system (100) of claim 1, wherein the predictive recognition and notification module includes a visualization interface that highlights the erroneous text and displays suggestions for neutral language alternatives directly in the place pronunciation.The system (100) of claim 1, wherein the user interface provides an option for users to upload place writings from multiple file formats including text files, PDFs, and word processing documents.The system (100) of claim 1, wherein the integration module enables automatic analysis of location prescriptions while created or edited within a recruiting software platform and provides sex-specific distortion feedback before the prescription is published.The system (100) of claim 1, wherein the system includes a learning mechanism that allows user feedback on marked pre-taken terms to be included in the training dataset so that the system may improve over time.