Global Journal of Computer Science and Technology | Volume XIII Issue III Version I
Year 2013 16
CAPTCHA: Attacks and Weaknesses against OCR Technology

ubiquitous defense used to protect open Web resources from being exploited at scale. An effective CAPTCHA resists existing mechanistic software solving, yet can be solved with high probability by a human being, producing more reliable human/computer distinguishers.

They have argued that CAPTCHAs, while traditionally viewed as a technological impediment to an attacker, should more properly be regarded as an economic one, as witnessed by a robust and mature CAPTCHA-solving industry which by passes the underlying technological issue completely.

Image to Text error rate for the custom Asirra CAPTCHA over time
Figure 2 : Image to Text error rate for the custom Asirra CAPTCHA over time [2]

CAPTCHAs are suitable for use with standard solver image APIs. The authors wrote the instructions “Find all cats” in English, Chinese (Simplified), Russian and Hindi across the top, as the majority of the workers speak one of these languages. They submitted this image once every three minutes to all services over 12 days.

Image to Text displayed a remarkable adaptability to this new CAPTCHA type, successfully solving the CAPTCHA on average 39.9% of the time. Figure 2 shows the declining error rate for Image to Text; as time progresses, the workers become increasingly adept at solving CAPTCHA.

Elie Burstein, et al, [7] describe that the CAPTCHAs are designed to be easy for humans but hard for machines. However, most recent research has focused only on making them hard for machines. In this paper, they presented what is to the best of their knowledge the first large scale evaluation of CAPTCHAs from the human perspective, with the goal of assessing how much friction CAPTCHAs present to the average user. For the purpose of this study they have asked workers from Amazon’s Mechanical Turk and an underground CAPTCHA breaking service to solve more than 318000 CAPTCHAs issued.

In the paper, “Text-based CAPTCHA Strengths and Weaknesses“ by Elie Burstein and Mathieu Martin [4], The authors carry out a systematic study of existing visual CAPTCHAs based on distorted characters that are augmented with anti-segmentation techniques.

Applying a systematic evaluation methodology to 15 current CAPTCHA schemes from popular web sites, the authors find that 13 are vulnerable to automated attacks. Based on this evaluation, we identify a series of recommendations for CAPTCHA designers and attackers, and possible future directions for

In the paper, “Attacks and Design of Image Recognition CAPTCHAs” [6], the authors systematically study the design of image recognition CAPTCHAs (IRCs). They first reviewed and examine all IRCs schemes known to them and evaluate each scheme against the practical requirements in CAPTCHA applications, particularly in large-scale real-life applications such as Gmail and Hotmail. Then the authors present a security analysis of the representative schemes the authors have identified. For the schemes that remain unbroken,

In the paper they presented their novel attacks. For the schemes for which known attacks are available, the authors propose a theoretical explanation why those schemes have failed. The authors have attempted a systematic study of image recognition CAPTCHAs.

The authors provided a thorough review of the state-of-the-art, presented a novel attack on a representative scheme, and analyzed successful attacks on the other representative schemes. Learned from these attacks, the authors defined for the first time a simple but novel framework for guiding the design of robust image recognition CAPTCHAs.

The framework led to their design of Cortcha, a novel CAPTCHA that exploits semantic contexts for image object recognition. Their usability study showed that Crotch yielded a slightly better human accuracy rate than Google’s text CAPTCHA. Cortcha offers the following novel features.

III. Conclusion

It is possible to enhance the security of an existing text CAPTCHA by systematically adding noise and distortion, and arranging characters more tightly. These measures, however, would also make the characters harder for humans to recognize, resulting in a higher error rate. There is a limit to the distortion and noise that humans can tolerate in a challenge of a text CAPTCHA. Usability is always an important issue in designing a CAPTCHA. With advances of segmentation and Optical Character Recognition (OCR) technologies, the capability gap between humans and bots in recognizing distorted and connected characters becomes increasingly smaller. This trend would likely render text CAPTCHAs eventually ineffective.

In the paper,” The Robustness of Google CAPTCHAs” [5], they reported a novel attack on two CAPTCHAs that have been widely deployed on the Internet, one being Google's home design and the other acquired by Google (i.e. re CAPTCHA). With a minor change, their attack program also works well on the latest Re CAPTCHA version, which uses a new defense mechanism that was unknown. When they designed their attack.

© 2013 Global Journals Inc. (US)