Plagiarism Detection in Spoken Responses Using ML Similarity Grids
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Plagiarism in spoken responses during language assessments, particularly in non-native speaking proficiency tests, is challenging to detect due to the availability of online resources, affecting the validity of scores and requiring efficient automated detection methods.
Innovation Solution
The use of machine learning models, specifically deep convolutional neural networks, to analyze similarity grids generated from transcribed spoken responses and source materials, incorporating techniques like exact word matching, stemming, and word embeddings to identify plagiarized content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human raters manually review spoken responses to detect plagiarism, then detection accuracy can be maintained at a high level, but the process becomes time-consuming and labor-intensive
Solution Approach 1:
The patent segments the plagiarism detection process into distinct automated stages: transcription of spoken responses, generation of similarity grids comparing transcribed text against source materials, and machine learning-based analysis of these grids. This segmentation enables parallel processing and eliminates the need for sequential manual review, thereby maintaining detection accuracy while dramatically reducing the time required.
Solution Approach 2:
The patent introduces an intermediary automated system comprising speech recognition software and machine learning models that act as a mediator between the spoken response and the plagiarism detection process. This intermediary automatically transcribes speech, generates similarity comparisons, and flags potential plagiarism, replacing the need for direct human review while preserving detection effectiveness.
2Productivity
If automated speech recognition and similarity comparison methods are used to detect plagiarism, then processing speed and efficiency are improved, but detection accuracy may deteriorate due to the complexity of identifying subtle plagiarized content
Solution Approach 1:
The patent transforms the plagiarism detection problem from direct text comparison into a different dimensional space by generating similarity grids that visualize relationships between transcribed responses and source materials. These grids convert complex textual similarity into a structured format that machine learning models can efficiently analyze, enabling both high-speed processing and accurate detection of plagiarized content including subtle modifications.
Solution Approach 2:
The patent changes the parameters of analysis by using machine learning models trained on similarity grid data rather than relying on simple string matching or manual review. The system adjusts parameters such as similarity thresholds and grid resolution to optimize both processing speed and detection accuracy, enabling the automated system to identify plagiarized content with precision comparable to human experts.
3Reliability
If comprehensive manual searching of source materials is conducted to verify plagiarism, then detection thoroughness is improved, but the complexity and resource requirements of the system increase
Solution Approach 1:
The patent creates a universal automated system that performs multiple functions: transcribing spoken responses, searching source materials, generating similarity comparisons, and detecting plagiarism. This multi-functional system consolidates what would otherwise require separate manual processes into a single integrated platform, maintaining thoroughness while reducing operational complexity and resource requirements.
Solution Approach 2:
The system enables self-service plagiarism detection by automatically performing source material searches and comparisons without requiring manual intervention. The automated speech recognition and similarity analysis tools independently identify and flag potential plagiarism, eliminating the need for human raters to manually search sources while maintaining comprehensive detection coverage.
Data Source
AI summary
Data is received that encapsulates a spoken response to a test question. Thereafter, the received data is transcribed into a string of words. The string of words is then compared with at least one source string so that a similarity grid representation of the comparison can be generated that characterizes a level of similarity between the string of words and the at least one source string. The grid representation is then scored using at least one machine learning model. The score indicates a likelihood of the spoken response having been plagiarized. Data providing the encapsulated score can then be provided. Related apparatus, systems, techniques and articles are also described.


