Preserving meaning over detector scores at HumanizeMy.ai

Fırat Mıhcı, Founder & Computational Linguist at HumanizeMy.ai
The AI Reality Check

Preserving meaning over detector scores at HumanizeMy.ai

Fırat Mıhcı is founder & computational linguist of HumanizeMy.ai, a research-driven service that analyses style signals to help writers and editors understand AI-detection limits. The company publishes methods and prioritises preserving an author’s meaning rather than chasing higher detector scores.

Fırat MıhcıFounder & Computational Linguist at HumanizeMy.ai
Published August 13, 2026
The takeaway

Use AI for routine, verifiable work where errors are observable and reversible, but keep a named human responsible for semantic approval and public claims; publish evidence and document judgment to build trust.

01

What’s the quick origin story of your business, and what makes what you do genuinely different from the other options your customers are looking at?

HumanizeMy.ai started with a question that had nothing to do with starting a software company: could you write a mathematically “correct” melody? I was interested in whether tension, surprise, repetition, and resolution could be measured. From there, I began asking whether good prose had a similar statistical shape.

My first hypothesis failed. Across three independent public benchmarks, the global “surprisal arc” of a document performed at chance; it could not reliably tell me who wrote something or whether it was good. That was frustrating, but useful. It stopped me looking for one magic number and pushed me toward the more complicated signals: regularity, self-similarity, register, coherence, and how those features interact.

That work became HumanizeMy.ai. What makes us different is that we do not treat an AI-detector score as ground truth, and we do not promise permanent “undetectability.” Our product sits on top of an ongoing computational-linguistics research program. We began with 2,590 real essays containing more than five million words. We later analysed 18,989 arXiv abstracts to see how AI-associated vocabulary was entering the human baseline, and tested 510 passages from five models answering the same 102 prompts. Style separated the model company surprisingly well, AUC 0.96, but identified the exact model only 50% of the time. That difference matters: a useful signal is still not a forensic verdict.

The unglamorous part is that I have thrown away approaches that looked great in a dashboard. Some rewrites improved the detector score but removed a qualification, added an implication, or changed the author’s claim. I did not ship them. If the meaning changes, the metric is irrelevant.

So our real distinction is not “we beat every detector.” It is that research sets the limits of what we claim, and preserving the writer’s meaning outranks producing an impressive score. We publish the methods, findings, and limitations through our research hub and ResearchGate so people can inspect the evidence themselves.

02

Where does most of your new business come from today, and where do you wish it came from?

HumanizeMy.ai is self-serve, so new business usually means people discovering the site and trying the product rather than us signing a fixed number of client contracts. Most of those people currently come through organic search.

They are rarely browsing casually. Usually something stressful has happened: a detector flagged an essay or article, an editor said the prose sounded mechanical, or the person tried several tools and got contradictory scores. They may first find one of our articles about false positives, look at the research behind it, and only then use the product. That is a slower route than an advertisement, but it tends to bring people who understand both what the tool can do and what no detector can prove.

Recommendations are the other route I value. Sometimes someone shares a result or a research page with a colleague facing the same problem. That kind of discovery is hard to manufacture, but it is meaningful because the trust comes from a real experience rather than a headline.

I would like more growth to come from the research itself: citations, educators, editors, writing professionals, responsible-AI teams, and eventually AI assistants that surface our evidence accurately. I would also like more customers to arrive because another user said, “It preserved what I meant and saved me a full manual rewrite.” That is a healthier signal than ranking for the fashionable phrase of the month.

I learned the weakness of depending too heavily on search when a major update exposed that some of our publishing had moved faster than its real differentiation. The tempting move was to publish even more. I did the opposite: paused expansion, audited the content, removed pages that did not justify their existence, and tightened the evidence and editorial gates.

I still want search, including AI-assisted search, to introduce people to HumanizeMy.ai. I just do not want the company to win by producing the largest pile of content. I want it to be found because the research or product solved a problem well enough to be cited, recommended, and remembered.

03

What’s one thing you’ve actually handed over to AI in the last year, and one thing you tried it on and went back to doing the human way?

I have handed AI a lot of the repetitive first-pass work inside our research and product process: drafting code scaffolds, suggesting edge cases, organising text batches, preparing feature-extraction steps, and turning an experimental plan into a reproducible checklist. It also helps with routine build and deployment work where we have tests and a rollback path.

That sounds less exciting than saying AI “runs the company,” but it is where the value is most real for me as a solo founder. It removes mechanical setup and gives me more time for the part that is difficult to automate: deciding whether the experiment is valid and what the result actually means. The tasks I delegate have something in common: the output can be checked against code, data, tests, or a written specification. If AI gets them wrong, the error is visible and reversible.

The thing I tried to hand over and took back was final semantic and editorial approval. During development, some rewriting approaches produced excellent detector numbers. On the dashboard, they looked ready to ship. But when I reviewed the outputs closely, some had removed a qualification, added an unsupported implication, or changed how strongly the author was making a claim. The prose was smoother and the score was better, but the meaning was not reliably intact. I stopped those approaches.

That led to a very simple rule: AI cannot approve its own work. It can propose, organise, compare, and flag. A person still decides whether a research conclusion is supported, whether a rewrite preserved meaning, and whether a public claim is honest. Before something ships, I want to know: did we omit or invent anything, does the conclusion outrun the evidence, and would I be comfortable explaining this result directly to the user?

I have not gone back to writing every line or running every step manually. The lesson was to delegate according to verifiability. AI gets the work where mistakes are cheap, observable, and reversible. Human judgment stays where a quiet mistake could change meaning or damage trust. That division has been much more useful than trying to automate everything.

04

Have you noticed a change in how customers find you or what they already know before they reach out? Is anyone arriving through AI assistants like ChatGPT yet?

Yes, the biggest change is that people arrive knowing much more of the vocabulary, but they are also more sceptical. A few years ago, the request was often simply, “Can you make this sound human?” Now people ask about false positives, model-specific style, perplexity, burstiness, second-language writing, or whether a detector score can really establish authorship. Some have already put the same text through several tools and want to know why every tool gave them a different answer.

That changes the conversation. I spend less time explaining that AI detection exists and more time explaining what a probabilistic result can and cannot prove. A score like 92% looks precise, but its meaning depends on the genre, population, model, and validation data behind it. In our own controlled work, stylistic features separated the model company much better than they identified the exact model. That is useful evidence, but it is not permission to make a forensic claim about who wrote a document.

AI assistants are becoming part of discovery too, although I would not pretend we have a perfectly clean number. Attribution is messy. A person may ask ChatGPT a question, open a cited page, compare products in Google, and then type our address directly. The analytics may only show the last step. We see early signs of AI-mediated discovery and receive questions that clearly seem shaped by an assistant’s explanation, but I treat that as directional evidence rather than a precise acquisition channel.

What has definitely changed is that discovery is becoming answer-first. Someone can form an opinion about HumanizeMy.ai before visiting the site. That is why we publish actual corpus sizes, methods, results, and limitations, and keep the company identity consistent across our site, ResearchGate, Google Scholar, and ORCID. If an assistant summarises our work, I want it to retrieve something inspectable rather than repeat a marketing superlative.

The upside is that customers can arrive much further into the decision. The downside is that they can arrive with a confident but incorrect summary. Being discoverable is no longer enough; the evidence has to be easy for both people and AI systems to represent accurately.

05

If you were advising someone in your industry on all of this for the next 12 months, what would you tell them to actually do, and what would you tell them to ignore?

For the next 12 months, I would start by writing down what AI is not allowed to overrule. At HumanizeMy.ai, preserving meaning outranks improving a detector score, and a public claim needs a dated method, real evidence, and an honest limitation. Those rules sometimes kill an attractive experiment. That is the point. They protect the customer when the numbers are tempting.

Then I would build a small evidence base of my own. Use independent data, publish the method, keep negative results, and make important analyses reproducible. Generic content is now almost free to produce. Original evidence is what gives customers, journalists, researchers, and AI assistants a reason to cite a business.

I would also manufacture disagreement. A small company cannot afford to let the same mind create an idea, represent the customer, judge the evidence, and approve the result without opposition. Use real user behaviour, independent benchmarks, specialist review, written pre-mortems, and small reversible tests. Feedback only counts if it is allowed to change the plan.

The least glamorous advice is to document repeated judgment, not just repeated tasks. If you have weighed the same risk three times, write the criteria. If every release needs factual, semantic, and rollback checks, make them gates. Save the reason an experiment failed. A negative result becomes valuable when it prevents you from paying for the same lesson twice.

For AI-assisted discovery, measure what you can without pretending attribution is clean. Track visible referrals, ask users how they found you, and check how assistants describe the brand. Make the identity, evidence, and limitations consistent wherever those systems might look.

I would ignore the race to publish the most AI-generated content, claims of perfect detection or permanent “undetectability,” screenshots presented as proof, and the idea that owning more AI tools makes a company advanced. I would also ignore pressure to automate every judgment. The most expensive failures often happen because nobody remains accountable for the result.

Use AI where errors are visible and reversible. Keep a named person responsible where meaning, evidence, or trust is at stake. Scale the quality of your decisions before you scale the volume of your output.

Thank you to Fırat Mıhcı and the team at HumanizeMy.ai for sharing what actually worked with Leaders Perception readers.

Share this interview

Want to share your playbook?

Leaders Perception publishes founder and operator interviews every week. A few questions, about five minutes, no calls. It is free and everyone we accept gets published.

Get featured on Leaders Perception

Explore additional categories

Explore Other Interviews