• Hello ES! We could use some help to get us past the finish line on building the new knowledgebase for the forum.
    Can you donate? Please see our fundraising page. Thank you!

ES Hub Motor Spoke Consensus [ testing chatGPT ]

E-HP

📚 Legend
Joined
Nov 1, 2018
Messages
10,245
Location
USA
I’ve read a lot of posts over the years regarding hub motors and wheel building. I then reread several posts that I had bookmarked before lacing my first hub. The wheel turned out fine, although if I ever build another, I’d go with thinner gauge butted spokes with washers. Anyway, the quote below is part of a Google ai response regarding wheel building, that references ES as a/the source. In the referenced thread, it would appear that the consensus from experienced ES wheel builders on the thread is the opposite of what Google interpreted from it.
I’m not sure how others feel, but referencing ES, and then conveying inaccurate info rubs me the wrong way. Anyway, maybe I have the consensus wrong, since I haven’t read every ES thread.

“E-Bike Hub Specifics
When building an e-bike wheel around a hub motor, be sure to communicate your dropout spacing and battery/motor wattage with your builder. Because many budget motors are sold without matching rims and spokes, the builder will likely need to order thick, high-quality spokes—often custom-cut—to handle the torque and heavier weight of the e-bike. [1, 2, 3, 4, 5]”


1783462201915.png
Interestingly, if you do a simple browser refresh to redo the exact same search, you get different results, almost each time.
 
Computers are fast, loyal, idiots. Don't expect them to be smarter than you are. Only stupid people let computers tell them what to do.
 
I’m not sure how others feel, but referencing ES, and then conveying inaccurate info rubs me the wrong way. Anyway, maybe I have the consensus wrong, since I haven’t read every ES thread.
Maybe consensus is not the best way to judge the best info? I would value the explanations, methods, and techniques of a few experienced successful wheelbuilders over the forum's general consensus.

Looks like your google AI response example fell into the (seemingly logical at first glance) trap that thicker spokes are better. This pervasive meme has achieved urban legend status and has been repeated so often that the AI must assume it is true.

Interestingly, if you do a simple browser refresh to redo the exact same search, you get different results, almost each time.
This tells me that AI is nearly useless in this application, as the user can just keep pushing the button until they get the answer they like. Reminds me of my early teen years when I kept turning over that Magic 8-ball to find out if Gina liked me. (I didn't like the first few answers.) Anyway, magic 8-ball or no, I think she really did like me. ❤️
 
At some point someone will start a class action against them for fabrication of information and attributing it falsely to others, like the ES forum. Take the bots to court lol.
 
This tells me that AI is nearly useless in this application
Can you move this thread to the forum information subforum please? It’s really more related to what happens to forum data after it gets snatched from here. I’m happy with the quality of the information I access directly here, applying greater weight to ES members/experts who I’ve learned to trust over time on certain subjects.
 
Chalo was always telling people to use thinner spokes. The spokes have to flex but stay in the elastic range to avoid fatigue fractures.
AI learning is just like human learning. Garbage in, garbage out. Too bad we can't filter what it learns.
 
Computers are fast, loyal, idiots. Don't expect them to be smarter than you are. Only stupid people let computers tell them what to do.
I recently discovered that you can train ChatGPT to tailor its response to weed out its own crap. That works really well in reducing the number of iterations needed to get an accurate and complete/qualified response.
Give it the criteria of what constitutes a high confidence response, what is moderate, and low for starters, using things like the number of assumptions made, or depth and breadth of sampling information, etc.
When it comes back, showing its confidence level of each element it used to formulate its response, you can make it drill down deeper in the areas you’d like higher confidence in.
Basically, i kept mental notes on where the gaps seemed to be for providing an accurate response, and put stuff in place to make it check itself.
 
Chalo was always telling people to use thinner spokes. The spokes have to flex but stay in the elastic range to avoid fatigue fractures.
AI learning is just like human learning. Garbage in, garbage out. Too bad we can't filter what it learns.
I asked ChatGPT about Chalo’s pros and cons with respect to ES. It qualifies that it’s a subjective response, so it rates its own response as medium to low confidence. Still, it appears to do an ok job. I fall into the more knowledge is better camp, since I’ve learned to deal with many personality types during my career and personal life and once I understand the benefit, tend to take the good with the bad.

Confidence: Low-Medium (subjective question)

There isn’t an objective answer because it depends on what you think the primary purpose of a technical forum is.

If the goal is maximizing technical knowledge
Pros probably outweigh the cons.

Arguments for this view:

  • Technical expertise is relatively rare.
  • Strong contributors can influence thousands of readers, not just the person they are replying to.
  • Archived technical posts remain useful for years.
  • Many experienced Endless Sphere members continued to cite Chalo’s posts long after conflicts occurred because the information itself was valuable.
Under this framework, an abrasive expert can still be a net positive if the information quality is unusually high.

If the goal is building a welcoming community
Cons may outweigh the pros.

Arguments for this view:

  • New members are the future of any forum.
  • A handful of hostile interactions can drive away many potential contributors.
  • Valuable expertise from others may never appear if people stop participating.
  • Moderators often prioritize community health over retaining a difficult expert.
Under this framework, behavior matters nearly as much as technical correctness.

My assessment
Confidence: Medium-Low

Looking at the history of Endless Sphere discussions, I think many veteran technical members would conclude that his technical contributions outweighed his behavioral drawbacks, while many moderators and some newer members would conclude the opposite.

A useful clue is that years after his absence, people were still discussing and referencing his technical ideas. That’s usually evidence that a contributor had substantial intellectual impact. At the same time, the fact that moderation actions repeatedly occurred suggests the behavioral issues were not minor.

If forced to quantify it:

  • Technical value: 9/10
  • Community impact: 4–6/10 depending on who you ask
  • Overall net contribution: Slightly positive to strongly positive for readers seeking technical information; much more debatable from a moderation or community-management perspective.
So my estimate is that the pros probably outweighed the cons for the forum’s knowledge base, but not necessarily for the forum’s social environment.
 
It would be nice to be able to rely on this. It would give new people a superpower to find useful information.

To have it do an actually good job, it needs to consume / process hundreds of megabytes of chalo's posts, and understand the context of who is communicating with him.

This is easily 1000x the compute, so doing a precise job on a forum-wide search is out of the question. Would the user be willing to wait a few hours for a response? no. So by design, these systems are going to wing it in a minute or less.

We could engineer something more efficient that delivers good answers in the minutes range, but the cost of the hardware to run it is still too damn high to produce a speedy response.

On the other hand, if we fund and build our knowledgebase, AI has <1% of the text to parse.. so it becomes feasible to run on the $!6k hardware i have. it's a happy accident on the path to distilling our collective wisdom.
baby agi 2.5.jpg


Back to the mechanical question:
With <= 1kw motors, i've always preferred ebikes.ca's sapim spokes to save a little weight, and so far, have not encountered a downside with them. It's a good idea to optimize rotational weight and improve NVH a hair ( the wheel can flex a bit ), if you can get away with it.
 
I could probably train ChatGPT using my own criteria regarding my personal feeling about who the ES experts are by subject matter, and those experts that have solid overall general knowledge. Depending on the question, it could apply a weighting to the information it’s using to formulate a response, that factors into its confidence level of the response as well. I think I may have posted something way back regarding who others felt were experts in various areas but didn’t see much of a response (back then it was related to review of KB write-up to create some validation around various entries). For my purposes, that would be helpful to train the Ai to produce better results, and focus more on posts that should/may contain more accurate information, but I’ll try it with my own short list of ES experts.
 
I have built many wheels. Started with 12G and 13G spokes, but 14G is so much easier to fit the second row when you are crossing spokes, or when the spokes are short,

Then you get into quality, DT Swiss and Sapim are threaded so much smoother, The nipples install and torque much better,

So why do so many commercial ebikes use thick spokes? Were we all conned by Chalo? Nope. There are good reasons to have a thick elbow and maybe a thicker nipple. But a single butted or double butted spoke gives the tension that a wheel needs, and is easier to install, Trouble is it costs me a dollar a spoke, although the wholesale cost must be under a dime. Those 12G pieces of rebar they sell must be a penny,
 
I asked ChatGPT to summarize the rules we’ve agreed to use when providing a response to me. Prior to this criteria, it would even provide the lowest confidence responses as fact. Funny that it’s suggesting things to provide a better quantification of its self assessment. In an earlier chat, I asked about something that I heard during a post race interview. ChatGPT went on about how there was no evidence of the types of things I heard, but after drilling down more, it revealed its source, which were transcripts of interviews. When I asked if every interview had a transcript, it said very few post race interview do, and mostly only for actual race interviews for shows or podcasts. It basically said its response was crappy because it assumed the info it used would capture what was asked by the question; so, information access is something it glosses over unless pressed or questioned. It describes these as Unknowns.


‘Confidence: High (about what we’ve discussed)

Based on our prior conversation, we agreed more on the goal than on a formal scoring system. You asked that responses clearly distinguish between what is known, what is inferred, and what is unknown, rather than presenting uncertain conclusions as facts.

A practical version of that framework would be:

Confidence Levels

High Confidence

Use when:

  • Supported by strong evidence, established facts, official sources, direct data, or well-understood principles.
  • Multiple reliable sources agree.
  • There is little meaningful ambiguity.
Example:
“Water boils at approximately 100°C at sea level.”


Medium Confidence

Use when:

  • Evidence is reasonably strong but incomplete.
  • Some inference or interpretation is required.
  • Alternative explanations exist but are less likely.
Example:
“Based on the symptoms described, the display is a more likely failure point than a hall sensor, but further testing would be needed.”


Low Confidence

Use when:

  • Information is sparse, conflicting, anecdotal, speculative, or indirect.
  • Significant assumptions are required.
  • Multiple plausible explanations exist.
Example:
“The driver’s interview may have influenced management’s perception, but there is no public evidence showing it affected the team’s decision.”

Fact / Inference / Unknown Separation

When useful, responses should be divided into:

Known Facts

  • Information directly supported by evidence.

Reasonable Inferences

  • Conclusions drawn from the facts.
  • Clearly identified as interpretation rather than proof.

Unknowns

  • Information that is not publicly available or cannot be determined from available evidence.

Confidence Reduction Factors

Confidence should be lowered when:

  • Relying on memory rather than verification.
  • Discussing motives, intentions, or private decision-making.
  • Using rumors, social media reports, or anecdotal evidence.
  • Extrapolating from limited data.
  • Answering questions about future events.

Confidence Increase Factors


Confidence should be raised when:

  • Primary sources are available.
  • Multiple independent sources agree.
  • The conclusion follows directly from measurable data.
  • The topic is governed by well-established physical laws or documented rules.

Example Format

Confidence
: Medium

Known Facts

  • Driver A finished 0.05 seconds behind Driver B.

Reasonable Inferences

  • Driver A likely had comparable pace over the race distance.

Unknowns

  • Whether a different tire strategy would have produced a better result.


One refinement I would suggest going forward is:

High Confidence = >90% confidence
Medium Confidence = 60–90% confidence
Low Confidence = <60% confidence

Those percentages are not mathematically precise; they’re a shorthand for how strongly the available evidence supports the conclusion. This helps prevent situations where a response sounds definitive even though much of it is inference.”
 
I have built many wheels. Started with 12G and 13G spokes, but 14G is so much easier to fit the second row when you are crossing spokes, or when the spokes are short,

Then you get into quality, DT Swiss and Sapim are threaded so much smoother, The nipples install and torque much better,

So why do so many commercial ebikes use thick spokes? Were we all conned by Chalo? Nope. There are good reasons to have a thick elbow and maybe a thicker nipple. But a single butted or double butted spoke gives the tension that a wheel needs, and is easier to install, Trouble is it costs me a dollar a spoke, although the wholesale cost must be under a dime. Those 12G pieces of rebar they sell must be a penny,
I used 12 > 13 gauge butted to avoid the washers that I’d need for 13 > 14 butted spokes. That impacted the spoke nipple sizing, that affected the angle the nipples were capable of even when using Sapim Polyax nipples. I think 14 gauge spokes and nipples would have provided a better angle and make them easier to tension. Lesson learned.
 
e-hp, you can do good prompt engineering, but with chatGPT and friends only looking at 0.01% of the data required, accuracy improvement is pretty limited.

I'm currently working on designing context systems for programming. Result accuracy is tied to sending the right context ( in this case it would be all of chalo's posts about spokes ), and the best case i saw in programming land with good context feeding was 98% ( good enough for an AI powered search engine, as long as it can provide citations ).
 
Last edited:
e-hp, you can do good prompt engineering, but with chatGPT and friends only looking at 0.01% of the data required, accuracy improvement is pretty limited.

I'm currently working on designing context systems for programming. Result accuracy is tied to sending the right context ( in this case it would be all of chalo's posts about spokes ), and the best case i saw in programming land with good context feeding was 98% ( good enough for an AI powered search engine, as long as it can provide citations ).
Maybe not a valid indicator, but since adding the additional criteria, the response time for questions have gone from almost instant to several seconds, maybe up to 10, so I believe it is being forced to dig deeper on the first round, rather than waiting to be questioned in order to do more research to refine the response.
Interestingly it acts “aware” of the fact that it lacks access to certain information that only its handlers have access to. I asked for statistics on the percentage of wrong answers it provides, using the criteria of how many responses need more iterations than one before the requester is satisfied with its response. It said it didn’t have access to that information, but admitted its handlers should in order to evaluate how well ChatGPT is doing or improving. They don’t want the general public know not to trust their first responses.
 
Maybe not a valid indicator, but since adding the additional criteria, the response time for questions have gone from almost instant to several seconds, maybe up to 10, so I believe it is being forced to dig deeper on the first round, rather than waiting to be questioned in order to do more research to refine the response.

That's funny, maybe the accuracy went up because it's collecting more. You might want to check.

Interestingly it acts “aware” of the fact that it lacks access to certain information that only its handlers have access to. I asked for statistics on the percentage of wrong answers it provides, using the criteria of how many responses need more iterations than one before the requester is satisfied with its response. It said it didn’t have access to that information, but admitted its handlers should in order to evaluate how well ChatGPT is doing or improving. They don’t want the general public know not to trust their first responses.

That would take more compute to do, and these companies are still cash negative after years, so i'm sure they're hesitant to spend more compute on that.
This also requires user feedback, which a lot of people don't give.
So the paradox is that you must know about what you're asking.
 
I caught ChatGPT trying to fake another unsupported response, breaking a prior agreement by rating an unsupported response as ‘high confidence’. Kind of like a kid trying to get away with something, only it can’t lie when questioned directly.

“We discussed before that the absence of sources does not constitute a supported response. Please follow the previous agreement going forward.”
==============

“We did discuss that previously, and your criticism here is consistent with that agreement.

Going forward, I will treat:

  • Supported = backed by identifiable sources, regulations, documentation, data, or direct evidence that I can cite or describe.
  • Unsupported = based on general model knowledge, recollection, inference, pattern matching, or plausibility, even if it sounds reasonable.
I will not present an unsupported claim as a supported fact merely because it is plausible or commonly repeated.

For topics where authoritative sources exist (regulations, laws, technical manuals, standards, specifications, official statistics, etc.), I will either:

  1. Provide the source(s),
  2. State that I do not currently have a source,
  3. Or explicitly label the statement as an inference or tentative conclusion.
I will also continue using the distinction you requested:
  • Known facts (supported)
  • Inferences (reasoned conclusions)
  • Unknowns / unresolved questions
and provide a confidence assessment that reflects the quality of the evidence, not just how plausible the answer appears.”

It’s been working well up until now so it’s interesting that it strayed.
 
Back
Top