<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[The new AI Captcha Module in BAS: full breakdown, setup, and real tests]]></title><description><![CDATA[<p dir="auto"><img src="/assets/uploads/files/1790979056286-00-thumbnail-ai-captcha-bas.png" alt="AI Captcha in BrowserAutomationStudio" class=" img-fluid img-markdown" /></p>
<p dir="auto">If you want, you can just watch the video on YouTube. I tried to fit both a tutorial and a lot of testing into one video, and I also used diagrams to make everything easier to understand</p>
<p dir="auto"><a href="https://youtu.be/qXkVmqoVlks" rel="nofollow ugc">https://youtu.be/qXkVmqoVlks</a></p>
<p dir="auto">I also created a telegram channel where I’ll share my thoughts and interesting stuff. For now, I’ve posted the mind maps from the video and some articles there: <a href="https://t.me/mustangbasen" rel="nofollow ugc">https://t.me/mustangbasen</a></p>
<p dir="auto">This article is basically a shorter text version. First we will look at how the module works and what is in the settings, then I will show what happened in <strong>real tests with different captchas</strong> and compare the models by results, requests, tokens and cost</p>
<p dir="auto"><img src="/assets/uploads/files/1790979080196-en-01-plan.png" alt="What we are going to look at" class=" img-fluid img-markdown" /></p>
<h3>Captchas I tested with real examples</h3>
<p dir="auto"><img src="/assets/uploads/files/1790979188080-en-06-tested-captchas.png" alt="CAPTCHAs used in the test" class=" img-fluid img-markdown" /></p>
<p dir="auto">I did not just look at the settings and assume the module worked. I actually ran it on different captchas and checked what it could solve. These were included in the tests:</p>
<ul>
<li>reCAPTCHA v2 <a href="https://youtu.be/qXkVmqoVlks?t=1582" rel="nofollow ugc">26:22</a></li>
<li>hCaptcha <a href="https://youtu.be/qXkVmqoVlks?t=2030" rel="nofollow ugc">33:50</a></li>
<li>GeeTest v4 <a href="https://youtu.be/qXkVmqoVlks?t=2691" rel="nofollow ugc">44:51</a></li>
<li>Cloudflare Turnstile <a href="https://youtu.be/qXkVmqoVlks?t=3049" rel="nofollow ugc">50:49</a></li>
<li>GeeTest Adaptive <a href="https://youtu.be/qXkVmqoVlks?t=3149" rel="nofollow ugc">52:29</a></li>
<li>GeeTest v3 <a href="https://youtu.be/qXkVmqoVlks?t=3417" rel="nofollow ugc">56:57</a></li>
<li>Basic text from an image <a href="https://youtu.be/qXkVmqoVlks?t=3835" rel="nofollow ugc">1:03:55</a></li>
<li>Click captcha <a href="https://youtu.be/qXkVmqoVlks?t=3916" rel="nofollow ugc">1:05:16</a></li>
<li>Rotate captcha <a href="https://youtu.be/qXkVmqoVlks?t=4111" rel="nofollow ugc">1:08:31</a></li>
<li>MTCaptcha <a href="https://youtu.be/qXkVmqoVlks?t=4333" rel="nofollow ugc">1:12:13</a></li>
</ul>
<p dir="auto">I also wanted to test Prosopo, but I could not test it properly because the test page was unavailable</p>
<h3>What AI Captcha actually does</h3>
<p dir="auto"><img src="/assets/uploads/files/1791028666028-89fb1a1a-63ce-4da7-8ec4-28c6eed79e2a-image.png" alt="How AI Captcha works" class=" img-fluid img-markdown" /></p>
<p dir="auto">Put simply, AI Captcha solves a visible captcha inside the built-in BAS browser. It takes a screenshot, sends it to the models, gets commands back and then clicks, drags something or types text into a field</p>
<p dir="auto">The model does not control the mouse by itself. It just looks at the image and tells BAS where to click, what to drag and where to type. BAS then does the actual work in the browser</p>
<p dir="auto">I tested all of this in the new <strong>BrowserAutomationStudio 30.9.0</strong> release</p>
<h3>Why there are two models</h3>
<p dir="auto"><img src="/assets/uploads/files/1790979080370-en-02-how-module-works.png" alt="How the two models work" class=" img-fluid img-markdown" /></p>
<p dir="auto">At this point you might be wondering why the module needs two models at all. It looks like BAS could just send the screenshot to one model and let it handle everything. Basically, the work is split so an expensive model is not used for every small check</p>
<ol>
<li>
<p dir="auto"><strong>The secondary model</strong> sees the whole BAS browser screen. It finds the captcha, works out its boundaries, can click a basic checkbox and then checks whether the captcha is solved. You can use something cheaper here because the job is usually pretty simple</p>
</li>
<li>
<p dir="auto"><strong>The main model</strong> receives the cropped captcha itself. This is the model that needs to understand what is on the image, what it is supposed to do and how to solve it. So this is where it makes sense to use a stronger model</p>
</li>
<li>
<p dir="auto">And <strong>BAS</strong> is basically the executor here. It takes screenshots, crops the captcha, checks the model response, maps the coordinates back to the page and moves the mouse</p>
</li>
</ol>
<hr />
<p dir="auto"><img src="/assets/uploads/files/1790979080122-en-03-model-actions.png" alt="Actions the model returns to BAS" class=" img-fluid img-markdown" /></p>
<p dir="auto">The main model can return three types of actions:</p>
<ul>
<li>
<p dir="auto"><code>click</code> clicks a point using the <code>x</code> and <code>y</code> coordinates</p>
</li>
<li>
<p dir="auto"><code>drag</code> moves something from the <code>from</code> point to the <code>to</code> point</p>
</li>
<li>
<p dir="auto"><code>type</code> clicks a field and enters text</p>
</li>
</ul>
<p dir="auto">It can also return several actions in one response. For example, it can select several images and then press the confirmation button. So BAS does not need a separate request for every single click</p>
<h3>What I entered in the BAS action</h3>
<p dir="auto"><img src="/assets/uploads/files/1790979031690-10-ai-captcha-action-main.png" alt="Main AI Captcha action settings" class=" img-fluid img-markdown" /></p>
<p dir="auto">I used <a href="https://openrouter.ai/" rel="nofollow ugc">OpenRouter</a> for these tests. The convenient part is that you can switch between different models through one provider. You just need to create an account there and copy your key, and I showed how to do all of that in the video</p>
<p dir="auto">In the video I go through OpenRouter in detail, starting with account registration: <a href="https://youtu.be/qXkVmqoVlks?t=994" rel="nofollow ugc">watch from 16:34</a></p>
<p dir="auto">Let's start with the main settings in the captcha-solving action:</p>
<ul>
<li>
<p dir="auto"><code>Main provider</code> is the provider</p>
</li>
<li>
<p dir="auto"><code>Main token</code> is the key for the main model</p>
</li>
<li>
<p dir="auto"><code>Main model</code> is the main model that will solve the captcha</p>
</li>
<li>
<p dir="auto"><code>Secondary provider</code>, <code>Secondary token</code> and <code>Secondary model</code> are the same settings for the secondary model that finds the captcha and then checks the result</p>
</li>
<li>
<p dir="auto"><code>Instructions</code> is where you can tell the model what kind of captcha you are solving and how it should solve it</p>
</li>
</ul>
<p dir="auto">I underestimated <code>Instructions</code> at first. A model can usually work out a simple captcha by itself. But when it needs to click objects in order, solve a puzzle captcha or understand some unclear mechanic, it starts messing up, wasting requests or saying it is done while the captcha is still sitting there. So, basically, try to write the instructions for the specific captcha you are solving right away</p>
<h5>How to add a model from OpenRouter</h5>
<p dir="auto"><img src="/assets/uploads/files/1791041495303-2884d7d5-b2ac-4e8c-ba75-ae3a85e28d3e-image.png" alt="How to move a model from OpenRouter to BAS" class=" img-fluid img-markdown" /></p>
<p dir="auto">If the model you need is missing from the ready-made list and you are using OpenRouter, open its page, copy the exact <code>Model ID</code> and paste the full value into the model field, just like I showed in the screenshot</p>
<p dir="auto">I chose three main models for this project:</p>
<ul>
<li>
<p dir="auto"><code>qwen/qwen3.8-27b:free</code> as the free one</p>
</li>
<li>
<p dir="auto"><code>google/gemini-3.8-flash</code> as the cheap one</p>
</li>
<li>
<p dir="auto"><code>anthropic/claude-sonnet-5</code> as the expensive one</p>
</li>
</ul>
<p dir="auto">Models and prices change all the time. There was no special logic behind this selection, I just wanted to test what each model could do and how the module worked overall</p>
<h4>Advanced settings</h4>
<p dir="auto"><img src="/assets/uploads/files/1790979031651-12-ai-captcha-action-advanced.png" alt="Advanced AI Captcha action settings" class=" img-fluid img-markdown" /></p>
<p dir="auto"><img src="/assets/uploads/files/1790979031666-13-ai-captcha-action-output.png" alt="AI Captcha action result and log" class=" img-fluid img-markdown" /></p>
<p dir="auto">There are a lot of settings in the Advanced tab, so here is what each one does:</p>
<ul>
<li>
<p dir="auto"><code>Main thinking</code> and <code>Secondary thinking</code> control how hard the models think about the task. Start with <code>low</code>. If the model completely fails, try <code>high</code>, because that really helped on many tasks</p>
</li>
<li>
<p dir="auto"><code>Max steps</code> is 14 by default. If the main model sends 14 responses and still gets nowhere, the action ends with an error. For some rotate captchas it is better to allow more steps</p>
</li>
<li>
<p dir="auto"><code>Check interval</code> and <code>Stability timeout</code> make BAS wait after a click or drag instead of taking another screenshot while the page is still changing</p>
</li>
<li>
<p dir="auto"><code>Timeout</code> limits the whole action by time. The default is 300 seconds (5 minutes)</p>
</li>
<li>
<p dir="auto"><code>API timeout</code> gives one model request up to 60 seconds. If you use free models, set it to 120 or even 180 seconds right away</p>
</li>
<li>
<p dir="auto"><code>API attempts</code> is 3 by default. That means one normal request and up to two more attempts if the provider temporarily fails</p>
</li>
<li>
<p dir="auto"><code>Main compression</code> and <code>Secondary compression</code> resize the images. More compression means fewer image tokens, but the model may simply stop seeing small text and details, so this needs testing</p>
</li>
<li>
<p dir="auto"><code>Similarity</code> helps BAS decide whether the page has stopped changing. The default is 97 percent</p>
</li>
<li>
<p dir="auto"><code>Summary</code> and <code>Full log</code> are just for debugging. <code>Summary</code> shows a short action history, while <code>Full log</code> saves the complete log with screenshots and page data</p>
</li>
<li>
<p dir="auto"><code>Result</code> is saved to <code>AI_CAPTCHA_RESULT</code>. It shows the number of steps and requests, token usage and runtime</p>
</li>
<li>
<p dir="auto"><code>Reason</code> is saved separately to <code>AI_CAPTCHA_REASON</code>. It contains the short completion reason, for example <code>model_done</code> or <code>no_captcha</code>. <code>model_done</code> still does not mean the site actually accepted the captcha, so it is better to check the page separately after the action</p>
</li>
</ul>
<h3>How the whole thing works step by step</h3>
<p dir="auto"><img src="/assets/uploads/files/1790979187836-en-04-solve-loop.png" alt="How one solve runs" class=" img-fluid img-markdown" /></p>
<p dir="auto">Basically, it works like this:</p>
<ol>
<li>
<p dir="auto">BAS takes a browser screenshot</p>
</li>
<li>
<p dir="auto">The secondary model checks whether a captcha is there and where it is</p>
</li>
<li>
<p dir="auto">If it is just a checkbox, the secondary model can click it itself</p>
</li>
<li>
<p dir="auto">If a challenge appears, BAS crops it and sends it to the main model</p>
</li>
<li>
<p dir="auto">The main model returns clicks, drags or text</p>
</li>
<li>
<p dir="auto">BAS performs the actions and waits until the page stops changing</p>
</li>
<li>
<p dir="auto">BAS takes another screenshot and checks everything again</p>
</li>
</ol>
<p dir="auto">This keeps going in a loop. It stops when the captcha disappears, the model thinks it is solved, the overall timeout runs out or the main model uses all 14 responses</p>
<h3>What kinds of captcha the module can solve</h3>
<p dir="auto"><img src="/assets/uploads/files/1790979187921-en-05-captcha-types.png" alt="Supported CAPTCHA types" class=" img-fluid img-markdown" /></p>
<p dir="auto">If a captcha can be solved from an image with clicks, dragging or text input, the module can at least try. That includes checkboxes, image selection, text from an image, sliders, image rotation, drag-and-drop and clicking objects in a particular order</p>
<p dir="auto">It obviously does not solve audio captchas or invisible captchas like recaptcha_v3</p>
<p dir="auto">I had some partial success with dynamic hCaptcha, but it never worked consistently. One challenge had a wasp flying between several flowers, and then you had to pick the flower it had never landed on</p>
<p dir="auto">That was where the model started messing up, because it receives separate images instead of a video. It simply could not properly track which flowers the wasp had landed on. Basically, maybe you could get it working by spending a lot of time on good instructions and testing different models, but challenges like this probably will not work right away, if they can even be solved with this module at all</p>
<p dir="auto"><img src="/assets/uploads/files/1791032627034-a79e71a5-42d4-4732-bb0d-863602510e2f-image.png" alt="Dynamic hCaptcha example" class=" img-fluid img-markdown" /></p>
<hr />
<h3>What happened with the models</h3>
<p dir="auto"><img src="/assets/uploads/files/1790979220487-en-07-model-comparison.png" alt="Main model comparison" class=" img-fluid img-markdown" /></p>
<p dir="auto">Looking only at the saved runs where the main model was actually called, I got this:</p>
<ul>
<li>
<p dir="auto"><strong>Qwen 3.8 27B free</strong> made 19 requests, used 49,952 tokens and cost $0</p>
</li>
<li>
<p dir="auto"><strong>Gemini 3.8 Flash</strong> made 44 requests, used 104,095 tokens and cost $0.097</p>
</li>
<li>
<p dir="auto"><strong>Claude Sonnet 5</strong> made 92 requests, used 306,285 tokens and cost $0.628</p>
</li>
</ul>
<p dir="auto">These costs include both the main and secondary models used in each run. I left Turnstile out because the main model was not called there at all</p>
<p dir="auto">And this is where things got interesting. Claude Sonnet 5 understood the image quickly and usually understood what it was supposed to do. But sometimes its coordinates were just awful. The much cheaper Gemini 3.8 Flash ended up working better on my GeeTest puzzles</p>
<p dir="auto">So yeah, choosing the most expensive model and expecting it to solve everything is not going to work. You need to take the actual captcha and see which model behaves normally on that particular task</p>
<h3>What happened with each captcha</h3>
<p dir="auto"><img src="/assets/uploads/files/1790979221008-en-08-results-by-captcha.png" alt="Results for each CAPTCHA" class=" img-fluid img-markdown" /></p>
<ul>
<li>
<p dir="auto"><strong>reCAPTCHA v2:</strong> there was not much to say here, just a normal image selection challenge. <strong>All three models solved it</strong></p>
</li>
<li>
<p dir="auto"><strong>hCaptcha:</strong> Gemini and Claude sometimes handled the static tasks, so they have a <strong>partial result</strong>. As soon as adaptive captchas appeared, everything started falling apart and the results were weak. There are a lot of different hCaptcha tasks, so I think the module can solve many of them, but the one with the wasp seems impossible to me)))</p>
</li>
<li>
<p dir="auto"><strong>GeeTest v4 and GeeTest v3:</strong> the free Qwen model completely failed. Claude seemed to understand the puzzle, but then returned terrible coordinates and moved the slider to the wrong place. <strong>Gemini solved the puzzle</strong>, and I ran GeeTest v4 again later and it solved it again</p>
</li>
<li>
<p dir="auto"><strong>Cloudflare Turnstile:</strong> this one was really simple. The secondary model clicked the checkbox and checked the result. <strong>The main model was not even needed</strong></p>
</li>
<li>
<p dir="auto"><strong>GeeTest Adaptive:</strong> Gemini solved it. Claude selected the correct objects but clicked slightly outside the right spots, so the whole thing failed. The free model first missed the captcha completely and then hit a free pool rate limit</p>
</li>
<li>
<p dir="auto"><strong>Text from an image and MTCaptcha:</strong> <strong>all three models solved them</strong>. There is not much to break down here, the model recognized the characters, entered them into the field and that was it</p>
</li>
<li>
<p dir="auto"><strong>Click captcha:</strong> Qwen would start correctly and then mess up later, so it only got a partial result. <strong>Gemini and Claude solved it in the saved runs</strong></p>
</li>
<li>
<p dir="auto"><strong>Rotate captcha:</strong> it basically ate more money than anything else. The free model got confused, Gemini hit the timeout and <strong>Claude finally solved it</strong>. Even Claude needed a ton of repeated checks, and the two saved runs averaged about <strong>$0.145</strong> each</p>
</li>
</ul>
<p dir="auto">There were 10 captcha types, 161 model requests and 474,965 tokens. The saved runs cost about <strong>~$0.727</strong>. These are not super accurate numbers, just a rough estimate. Overall I spent $3 on the whole project, which included 80-100 different captchas across three models, because a lot happened off camera and I also had to reshoot parts</p>
<h3>What I learned from all of this</h3>
<p dir="auto"><img src="/assets/uploads/files/1790979221131-en-09-test-findings.png" alt="What the tests showed" class=" img-fluid img-markdown" /></p>
<p dir="auto"><strong>The most expensive model is not automatically the best one.</strong> Claude understood the tasks faster, but its coordinates were awful and that broke the whole solve. Gemini was cheaper and simply worked better on several puzzles</p>
<p dir="auto"><strong>Good instructions make a real difference too.</strong> If you throw some weird captcha at a model without explaining anything, it may sit there thinking, mix up the order and burn through requests. I would explain what needs to be selected, where it needs to be dragged and when the model should press confirm, <strong>literally explain everything like you would to a small child</strong></p>
<p dir="auto">With <strong>thinking</strong>, I would keep it simple. Start with <code>low</code> and see whether the model can handle the captcha. If it keeps messing up, try <code>high</code>. There is not much point in using <code>high</code> everywhere from the start because it can mean <strong>more time and more tokens</strong></p>
<p dir="auto">And <strong>one AI Captcha action probably will not be enough in a real project</strong>. I saw provider errors, <code>invalid actions</code>, bad coordinates and models deciding they were done too early. So you need <strong>an error handler, a limited number of retries and a normal page check after the captcha</strong></p>
<p dir="auto"><strong>Free models can work too</strong>, but you should expect some hassle. They solved simple text captchas, reCAPTCHA and Turnstile in my tests. The shared <a href="https://openrouter.ai/" rel="nofollow ugc">OpenRouter</a> free pool was sometimes overloaded though, so I got a <strong>429 error</strong> and a simple captcha could take several minutes. It is fine if you just want to try the module, but for regular work I would look for something more stable</p>
<h3>So what is the result</h3>
<p dir="auto">Overall, <strong>I liked the module</strong>. It is not tied to one traditional captcha-solving service, and you can try different models to see which one works better for a particular captcha</p>
<p dir="auto">You need to understand that <strong>this is not one button where you click once and it solves everything by itself</strong>. You really have to <strong>sit down and test it</strong>: try different models, change the instructions, switch between <code>low</code> and <code>high</code> thinking, see where things break and run it again. Most of the result comes down to testing, because one model may solve one captcha normally and completely fail on another</p>
<p dir="auto">Everything about the action itself and its settings was checked in BrowserAutomationStudio 30.9.0. The models, prices and free-pool conditions were current when I ran the tests</p>
]]></description><link>http://community.bablosoft.com/topic/32566/the-new-ai-captcha-module-in-bas-full-breakdown-setup-and-real-tests</link><generator>RSS for Node</generator><lastBuildDate>Tue, 06 Oct 2026 02:39:32 GMT</lastBuildDate><atom:link href="http://community.bablosoft.com/topic/32566.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 05 Oct 2026 08:43:00 GMT</pubDate><ttl>60</ttl></channel></rss>